OpenAI and Ironclad have built a research evaluation for AI agents that configure complex contracting workflows inside specialised business software. Announced on 6 October, the collaboration turns legal, commercial and procurement requirements into measurable computer-use tasks that test whether an agent can preserve business rules across many steps.

GPT-6 Astra is the first OpenAI frontier model trained on the Ironclad tasks. On the research evaluation, its average score was 32 per cent higher than GPT-5.6 Sol’s, while estimated time per attempt was 48 per cent lower. OpenAI cautions that these are simulated estimates, not measured customer savings.

Eleven tasks reflect real contracting work

Ironclad staff and users at OpenAI identified 11 tasks, including configuring nondisclosure agreements, building procurement approvals and updating reusable clauses according to jurisdiction. An experienced user would take about 30 to 40 minutes per task, according to the estimate.

Each task was evaluated against eight to 50 criteria. That granularity reveals whether an agent completed all required paths rather than merely producing a plausible final screen. Ironclad provided hosted environments where models could practise and be evaluated safely.

Business rules must survive the full workflow

A software-purchase process may require Finance approval above a threshold, Security review for particular requests and Legal review for nonstandard terms. The agent must translate those requirements into forms, templates, approval logic and final records, then confirm that cases above and below each threshold behave correctly.

This is harder than executing isolated clicks. A workflow can appear complete while omitting an exception or routing one case incorrectly. Evaluations therefore need end-to-end tests that exercise the configuration under several conditions.

Astra improved both score and simulated speed

OpenAI reports that Astra satisfied more criteria with an estimated average attempt time of 19.2 minutes, compared with 37 minutes for GPT-5.6 Sol. In one example, Astra met about 94 per cent of the task criteria in 20 minutes, while the earlier model met about 85 per cent in 32 minutes.

The numbers cover only the 11 research tasks. They do not establish performance across every Ironclad workflow, and the time calculation is based on assumed processing and generation speeds. Production trials would need to include review, corrections and the consequences of a misconfigured contract process.

Domain experts define success

Ironclad’s contribution was not only software access. Its specialists helped identify high-value failures and convert tacit workflow knowledge into explicit criteria. This model of collaboration can make agent evaluations more realistic than generic browser benchmarks.

It also creates governance questions about data and representativeness. OpenAI says simulated contracts came from public SEC EDGAR material after filters to remove personal information, and that it did not use customer or nonpublic Ironclad contracts for training or evaluation.

Human oversight remains essential

Contracting systems encode authority, risk and legal obligations. Even a high evaluation score leaves room for a missed rule that could approve the wrong purchase or apply an unsuitable term. Human review, test cases and controlled publication of workflow changes remain necessary.

Organisations should log the agent’s actions and preserve the original requirements. Reviewers need to compare the implemented process with those requirements, not merely inspect whether the agent reported success.

OpenAI seeks more software partners

The company is inviting a small number of software vendors to propose difficult professional tasks, evidence of current model failure and criteria for success. Partners need domain experts, a secure test environment and data that can be used safely for research.

This approach could improve agents in specialised applications where public benchmarks offer little guidance. The risk is optimising too narrowly for selected partner environments, so evaluations should include novel variations and independent checks.

From demonstration to dependable work

The Ironclad study shows progress in sustained computer use, but its strongest contribution is methodological. Complex work can be decomposed into criteria without reducing it to one final screenshot, and domain experts can expose failures that model developers might overlook.

Before deployment, customers should reproduce the evaluation on their own rules and exceptions. A capable agent may reduce configuration effort, but accountability for the contracting process remains with the legal and business teams that approve it.

Production evidence should go beyond task scores

A real deployment should measure the frequency and severity of corrections, reviewer time and whether employees detect subtle configuration errors. It should also test changes in policy after the agent has learned an earlier pattern, because contracting rules are not static.

Safe adoption can use a staged model: the agent first observes, then proposes configurations in a sandbox, and only later publishes changes after approval. That approach creates evidence about reliability while keeping legal and financial consequences under human control.

Vendors should publish failure categories as well as aggregate scores. Missing an optional label is different from bypassing an approval threshold. Severity-weighted reporting would help customers understand whether improvements address cosmetic mistakes or the business rules that matter most.