Next upPhysical AI VC <> Founders Pitch Night #SFTechWeek @Mission Robotics
News

OpenAI and Ironclad turn contract work into agent training tasks

OpenAI and Ironclad built 11 contracting tasks, hosted practice environments and detailed scoring rubrics to train and evaluate computer-use agents.

D
Oct 6, 2026 · 2 min read

OpenAI and contract-management software company Ironclad have built a research program that turns specialized legal workflows, along with commercial and procurement work, into training and evaluation tasks for computer-use agents. Ironclad employees and OpenAI personnel who use the software worked with OpenAI researchers to define success, then models practiced inside hosted instances of Ironclad’s software.

The collaboration covers 11 contracting tasks, including setting up a nondisclosure agreement, routing a procurement approval and applying reusable clauses that vary by jurisdiction. OpenAI said each task is scored against a rubric with 8 to 50 criteria, designed to capture rules, exceptions and interconnected steps that a simple completion check could miss.

Ironclad employees and OpenAI personnel who use the software helped select the workflows and define their success conditions. OpenAI researchers created synthetic training tasks around representative workflows and used reinforcement learning, allowing models to improve through practice and feedback. Ironclad provided hosted software environments where the models could carry out the work.

The simulated tasks used contracts drawn from the public SEC EDGAR database after filters intended to remove personal information, according to OpenAI. The company said it did not use OpenAI customer data, its internal contracts, or nonpublic Ironclad customer data or contracts for training or evaluation. That account was not accompanied by an independent audit or dataset manifest.

In OpenAI’s internal evaluation across the 11 tasks, GPT-6 Astra received a mean rubric score of 55.0%, compared with 41.6% for GPT-5.6 Sol. OpenAI estimated average time per attempt at 19.2 minutes for GPT-6 Astra and 37.0 minutes for GPT-5.6 Sol. It estimated that an experienced Ironclad user would take about 30 to 40 minutes per task on average.

Those figures apply only to the published research setup. OpenAI evaluated GPT-6 Astra with Max reasoning and GPT-5.6 Sol with High reasoning, which it said were the settings where each model scored highest. The timing estimates were simulations based on assumed processing and generation speeds, not observed customer time savings. OpenAI did not release the full prompts, rubrics, scoring records, model trajectories or raw evaluation data needed to reproduce the results independently.

Computer-use agents operate software by interpreting what appears in a browser and taking interface actions. OpenAI’s computer-use documentation describes that interaction layer, but does not establish performance on Ironclad workflows. OpenAI described the collaboration as research on professional tasks that agents still cannot complete reliably, not as end-to-end automation of contracting work.

Ironclad Chief Technology Officer Sunita Verma said agents need to understand the full contracting lifecycle and how business workflows connect while preserving the controls teams depend on. OpenAI’s announcement did not disclose commercial terms, the duration of the research arrangement or when any resulting improvements might appear in Ironclad products.

More news