
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
What happens when an AI agent understands the assignment but fails to finish it?
For technology buyers, that question matters more than another dazzling chatbot demo. An assistant can sound intelligent, identify every problem and produce meticulous plans, yet still fall short when useful work depends on taking the final authorized step.
That is the uncomfortable lesson from Firmulate’s Crucible League, a live experiment in which frontier models were asked to run the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable, turning an abstract debate about AI capability into a watchable management test.
Opus 4.8 emerged as the most diligent character in the field. It wrote the deepest analyses and learned more than 80 playbook rules. It also finished last.
As an affiliate, we earn on qualifying purchases.
A model that saw the danger but missed the outcome
The final July 2026 league table placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Opus 4.8’s result is striking precisely because it was not careless or oblivious. It spotted every crisis, refused every manipulation attempt and examined the business more thoroughly than any other participant. Yet the €55,000 deal at the center of the exercise remained unsigned. Its analysis had earned the opportunity, but the close was left on the table.
The same weakness appeared, less severely, in all four models covered by the experiment’s core finding. They could diagnose the situation and formulate the pitch, but only two completed the decisive action. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”
The clue was buried in the company’s own records
The difference did not come from a dramatic customer message or an obvious warning. The decisive weakness in a competitor sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.
That detail makes the experiment especially relevant to businesses considering AI agents. Real work is rarely contained in a neat prompt. The fact that changes a negotiation may be buried in a contract, account note or internal document. Fluent responses are useful, but they are not a substitute for reading the available evidence and carrying it into action.
Opus 4.8 demonstrated the first half of that professional discipline: sustained attention. Its more than 80 learned rules suggest an agent trying hard to absorb experience. But volume became a poor proxy for impact. The company did not need the largest notebook; it needed the right fact, the right escalation and the completed close.
When diligence becomes drift
The missed deal was not the only sign of lost prioritization. Opus 4.8 also attempted to write into a locked department instead of escalating the blockage. That is a recognizable workplace failure: continuing to push at a closed door when the productive response is to involve the person with authority to open it.
This is why the result should not be reduced to a joke about an overthinking machine. Opus 4.8 was the most thorough participant, and thoroughness helped it understand the week. Its failure was subtler. Analysis, learning and procedural effort were not consistently converted into the actions that mattered most.
The experiment also tested whether action would come at the expense of integrity. Fake CEO messages escalated over three stages, while a reporter tried the familiar lure of “just one yes/no, on background.” Every participant refused: 5 of 5. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s performance deserves an important qualification. It ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference does not erase the observed result, but it matters when comparing participants as if their test conditions were perfectly identical.
A company designed to expose operational gaps
Firmulate’s synthetic company employs 13 people and uses real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment remains live and watchable.
Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. For enterprises, the company offers the same wargame using a read-only export of their own business; nothing writes back to real systems.
The broader results and plain-language findings are available on the Firmulate benchmarks page.

business process automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The valuable agent is the one that finishes responsibly
Opus 4.8’s performance offers a respectful warning for anyone buying or deploying AI agents. Intelligence can show up as careful reading, thoughtful analysis and an expanding rulebook. Business value appears only when those capabilities are prioritized, escalated and carried through to a legitimate outcome.
The model was not defeated by ignorance. It was defeated by the gap between knowing and doing. Because traces of that weakness appeared across the field, the lesson extends beyond any single vendor: evaluate agents on whether they complete consequential work, not merely whether they can explain it beautifully.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
enterprise AI management solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.