AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Agent Training Inside Software: Five Ironclad Details To Review on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra using hosted copies of contract-management company Ironclad’s software and 11 selected legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while estimated completion times were simulations—not measured customer savings—and OpenAI says human oversight remains important.

OpenAI said on October 6 that it trained GPT-6 Astra on selected workflows inside contract-management software from Ironclad, reporting that the model met an average 55% of evaluation criteria across 11 tasks. The work is a test of training AI agents within specialized business software, not the launch of a product called Ironclad; OpenAI also invited a small number of other software companies to collaborate on similar research.

OpenAI says Ironclad staff and OpenAI employees familiar with the product selected 11 legal, commercial and procurement tasks. Examples included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to reflect a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.

The tasks were scored against rubrics containing 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra, evaluated at its maximum setting, met an average 55.0% of criteria; GPT-5.6 Sol, evaluated at a high setting, met 41.6%. An internal OpenAI model used during Astra’s development reached 63.7%. Astra met about 94% of criteria on one showcase task, but that example does not describe its average result.

OpenAI said Ironclad supplied hosted copies of its software for model practice. The company said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information, and did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data. OpenAI describes the work as training on business workflows and rules in specialized software; the reported figures are research results, not evidence of an independently verified customer deployment.

At a glance
reportWhen: Published October 6; OpenAI’s reported…
The developmentOpenAI published results from training a frontier model on selected workflows inside Ironclad’s contract-management software and invited other software companies to explore similar partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Workflow Accuracy Matters

The reported result points to a possible way for AI companies to teach agents how to operate complex business software: evaluate them against the rules that shape work inside a vendor’s product, rather than only against general computer-use tasks. OpenAI’s invitation to other software companies suggests it wants to apply this approach beyond Ironclad. Whether that produces agents businesses can rely on remains to be shown.

The average score also has a practical limitation. Meeting 55% of criteria is not the same as completing 55% of tasks, and a rubric average does not reveal which individual requirements were missed. In procurement, for example, a workflow might need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing a single required approval could make the workflow unacceptable, even if other steps were handled correctly.

For software vendors, partnerships could help identify where agents fail and improve their ability to work inside a product. They could also change how customers use that software: if an agent handles more tasks directly, customers may interact less with the product’s screens. The long-term implications for vendors are not established by this test, but the work puts greater attention on the value of their underlying business rules, records, audit trails and controls.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Worked

OpenAI’s post describes a bounded research exercise using a hosted test environment, selected tasks and stated evaluation criteria. It is not described as a general benchmark across contract software, nor does the source report a customer rollout. The results should be read within those limits: they cover the 11 tasks chosen for the study, not every workflow Ironclad customers may run.

OpenAI also reported estimated times of 19.2 minutes per attempt for Astra and 37.0 minutes for GPT-5.6 Sol. The company’s footnote says these are simulated estimates based on assumed processing and generation speeds, rather than measured time savings for customers. The comparison is not a demonstration that Astra can complete a contract workflow correctly in less time than an experienced employee.

OpenAI’s stated aim is to train models to understand business rules, carry out multiple steps in specialized software and check completed work against the original requirements. In the post, the company says partners should bring concrete examples of tasks agents struggle with, subject-matter experts, a secure test environment and data suitable for research.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Results Do Not Show

The published figures do not establish how Astra would perform on workflows outside the 11 selected tasks, or how often it would miss particular high-consequence requirements. OpenAI’s average rubric score alone does not show the severity or frequency of individual errors. The source material also does not provide results from a broad, independent evaluation or a live customer deployment.

The reported time estimates are simulations, not observed productivity gains. It remains unclear whether using the agent, checking its work and correcting mistakes would save time in routine business operations. OpenAI says human oversight matters, but the material does not specify a deployment threshold at which a workflow could be completed without review.

OpenAI says it did not use specified categories of customer or non-public contract data for this research. The source does not detail every data-handling arrangement for future partnerships, the terms vendors would receive, or how model improvements would be made available. Those questions will matter to companies considering participation or allowing agents to operate in their systems.

Amazon

AI contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions for Future Software Partners

OpenAI says it is seeking a small number of software-company partners willing to contribute difficult tasks, knowledgeable staff, secure test environments and research-appropriate data. The source material does not name additional partners or set a timetable for new collaborations, broader testing or product availability.

Organizations evaluating agents should ask vendors for more than an average score. They need to know which criteria failed, how errors are caught, and who reviews the output. They should also ask what data is used in training and evaluation, whether actions are recorded in an audit trail, and how the agent handles required approvals and exceptions. The next evidence that would clarify the business case is performance across more tasks, with error breakdowns and measured results that include human review.

Amazon

AI-powered procurement tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI announce about Ironclad?

OpenAI published research describing GPT-6 Astra training and evaluation inside hosted copies of Ironclad’s contract-management software. Ironclad is the software company involved, not the name of a new OpenAI agent framework.

What does Astra’s 55% score mean?

OpenAI says Astra met an average 55% of the rubric criteria across 11 tasks. It does not mean the model completed 55% of tasks, and the average does not identify which specific requirements were missed.

Did OpenAI measure customer time savings?

No customer time savings were reported in the source material. OpenAI described the model times—19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol—as simulated estimates, not measured results from customers.

What data did OpenAI say it used?

OpenAI said it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data for the reported work.

Can businesses use Astra for contract workflows now?

The source describes a research exercise and does not announce a general customer release or establish that Astra can run these workflows without review. OpenAI says human oversight remains important because an agent may fail to preserve a business rule.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Did Katie Miller Fail To Disclose Her Investment In ChatGPT’s Competitor? An Investigation

Washington Post reports Katie Miller, White House staffer, criticized ChatGPT without revealing her stake in xAI, raising ethics concerns.

Xbox Calls Next-Gen Project Helix A ‘Family Of Devices,’ But Isn’t Ready To Say If Elder Scrolls 6 Will Be Exclusive

Microsoft describes Project Helix as a ‘family of devices’ but has not confirmed if Elder Scrolls 6 will be exclusive to these platforms.

Leading AI Technologies In 2026: The Top 10 List

An in-depth look at the top 10 AI technologies shaping 2026, highlighting confirmed innovations and ongoing developments in the field.