A partnership aimed at long procedures, not single answers

Consider what a legal operations team actually does when a company wants to buy software. Finance may need to approve any purchase above a certain dollar amount. Security may need to review certain requests. Legal may need to look at any nonstandard contract terms. The employee setting up the process has to turn that short list of rules into an intake form, document templates, approval rules, and a record of the final agreement. Getting one step right is not enough. The finished process has to work across every situation it was designed to handle, including the exceptions.

That kind of work, not a single clever answer, is what OpenAI says it is now training its frontier agent on. In a company publication dated October 6, 2026, OpenAI announced a research collaboration with Ironclad, a contract-software company, and reported results from its own research evaluation of GPT-6 Astra, its latest model, on eleven contracting and procurement workflows. The announcement is OpenAI's description of its own research. Every number in it is vendor-reported, and OpenAI itself footnotes that the headline time figures are simulated estimates, not measured customer time savings.

This article explains, for readers without a technical background, what the collaboration actually built, what the reported numbers do and do not show, and why the pattern behind it, teaching agents multi-step professional software by borrowing expert workflows, may matter beyond contracting.

Why contract software was a good first test

Many jobs are not a sequence of isolated questions. They are procedures with rules: approve purchases above a threshold, route nonstandard terms to Legal, make sure the form matches the policy. Software that encodes such procedures is everywhere in business, and using it well requires an agent to keep the whole rule set in view while it clicks through screens, rather than losing track halfway through.

In the announcement, OpenAI describes that exact failure mode. If an agent loses track of one of the business rules partway through a task, that limits what a software company can confidently ask it to do. Ironclad's role reflects a challenge its customers know well: a contracting process must handle exceptions while preserving the rules a business depends on.

This is the central idea of the collaboration: treat real professional workflows as research problems. Ironclad employees, plus people who use Ironclad at OpenAI, helped OpenAI's researchers identify eleven tasks spanning legal, commercial, and procurement work. Ironclad also provided hosted environments of its own product where the models could practice. The researchers then built synthetic training tasks and used reinforcement learning, a method in which models improve by practicing and receiving feedback, to train on them.

What was built: eleven tasks, 8 to 50 criteria each

OpenAI says the eleven tasks included setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause so it reflects the jurisdiction a requester selects. Each task was judged against between 8 and 50 criteria depending on complexity, which let OpenAI see which parts a model got right and where it fell short. The company estimates an experienced human user would take about 30 to 40 minutes per task on average.

The synthetic training tasks were built, according to OpenAI's footnote, from contracts publicly available in the SEC's EDGAR database after applying filters designed to remove personal information. OpenAI states that no OpenAI customer data, no OpenAI internal contracts, and no nonpublic Ironclad customer data or contracts were used for training or evaluation. That data-provenance point matters: it means the training material was public filings, not confidential business documents.

GPT-6 Astra, according to OpenAI, is its first frontier model trained on Ironclad tasks. That makes this collaboration a template, not a one-off: a software company contributes expert knowledge of what success looks like, hosted environments, and task definitions, and the model maker turns them into training and evaluation problems.

The numbers, and the footnote under them

OpenAI compared GPT-6 Astra, using its Max reasoning setting, against GPT-5.6 Sol, using its High reasoning setting, described as the settings where each model scored highest. On the eleven research tasks, OpenAI reports a mean rubric score of 55.0 percent for Astra versus 41.6 percent for Sol. That is the source of the announcement's claim that Astra's average score was 32 percent higher; note that is a relative comparison of two percentages, not a claim that Astra scores 32 points higher.

On time, OpenAI reports an estimated average of 19.2 minutes per attempt for Astra versus 37.0 minutes for Sol, described as 48 percent lower. An internal model used during Astra's development reportedly scored 63.7 percent on the same tasks, and OpenAI says it aims to bring those further gains to future models.

In one example shown in the release, Astra met about 94 percent of the task criteria in an estimated 20 minutes, while GPT-5.6 Sol met about 85 percent in an estimated 32 minutes.

All of these figures are vendor-reported results from OpenAI's own research evaluation. There is no independent verification of the rubric scoring, the task selection, or the time estimates. This is my assessment, not a fact from the source: self-reported evaluations by the company selling the model are a reasonable starting point but should be weighted as marketing-adjacent research until reproduced elsewhere.

What the release itself says it does not prove

OpenAI's release is unusually explicit about the limits. Footnote 1 states that the results cover the eleven research tasks, not all Ironclad workflows, and that the times are simulated estimates based on assumed model processing and generation speeds, not measured customer time savings. In other words, the minutes are not stopwatch measurements of anyone's workday; they are model-derived projections under assumptions.

That distinction matters for any reader who sees the numbers as a promise of productivity gains. What has been measured is performance on a rubric OpenAI designed with Ironclad's help, on tasks OpenAI and Ironclad selected. What has not been shown is how much time or money real customers would save, how the models perform on workflows outside the eleven tasks, or how performance would hold up on live customer data rather than synthetic and hosted environments.

Both companies, per the release, emphasize that human oversight still matters as agents improve at complex contracting tasks, and that a full contracting platform remains essential. That is also, fairly noted, a position that protects both companies' products; OpenAI sells models, Ironclad sells the platform the models would work within. Neither company claims the agent can replace the platform.

What Ironclad's CTO said

Sunita Verma, Chief Technology Officer of Ironclad, was quoted in OpenAI's release, which was published to announce the partnership and its reported results. In the release she stated:

For AI to be genuinely useful in complex areas of contracting, it has to do more than complete individual actions. Agents need to understand the full contracting lifecycle, including how business workflows connect while preserving the controls teams rely on. Our collaboration with OpenAI brings these real-world challenges into the research process and helps move the technology forward.

Attribution: Sunita Verma, Chief Technology Officer, Ironclad, quoted in OpenAI's company publication "Advancing computer use with Ironclad," October 6, 2026. The quotation is reproduced word-for-word from the retrieved primary source. It proves what she said, not that the collaboration will achieve those aims; assessing that requires the independent evidence that does not yet exist.

Open door: more software partners wanted

OpenAI says it is inviting a small number of additional software companies into the program through a research-collaboration application. The stated bar is concrete: applicants should bring an example of a task agents cannot reliably complete, evidence of where the agent fails, and a clear way of judging success, plus people who know the work deeply, a secure test environment, and data that can be safely used for research.

For the general reader, the pattern to watch is this: AI progress here is being driven less by abstract benchmarks and more by partners who define what a hard professional task looks like and how to grade it. If the approach generalizes, similar collaborations could emerge for other procedure-heavy professions, from accounting close processes to claims handling. Whether they deliver depends on exactly the things OpenAI's own footnotes flag: whether rubric success translates into real-world reliability, and whether the benchmarks survive independent scrutiny.

Facts, uncertainty, and opinion, separated

The verified facts are: OpenAI and Ironclad announced a research collaboration on October 6, 2026; eleven tasks were defined across legal, commercial, and procurement work, judged against 8 to 50 criteria each, with an estimated 30 to 40 minutes of experienced-user time per task; OpenAI reports Astra scoring 55.0 percent versus 41.6 percent for GPT-5.6 Sol, with estimated attempt times of 19.2 versus 37.0 minutes, and an internal development model at 63.7 percent; training used synthetic tasks from public SEC EDGAR contracts with personal information filtered; OpenAI is soliciting more software partners.

The uncertainty is: all performance figures come from OpenAI's own evaluation; times are simulated; the tasks do not cover all Ironclad workflows; no independent party has reviewed or reproduced the results.

My opinion, clearly separated from the reported facts: this is a genuinely sensible method for teaching agents long, rule-bound professional procedures, and the openness about simulated times is to OpenAI's credit. But until an outside evaluator can run the same tasks, the right reading is that OpenAI has demonstrated progress on tasks it designed, which is evidence of a method, not yet proof of real-world impact.