AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

What happens when an AI agent understands the assignment but fails to finish it?

For technology buyers, that question matters more than another dazzling chatbot demo. An assistant can sound intelligent, identify every problem and produce meticulous plans, yet still fall short when useful work depends on taking the final authorized step.

That is the uncomfortable lesson from Firmulate’s Crucible League, a live experiment in which frontier models were asked to run the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable, turning an abstract debate about AI capability into a watchable management test.

Opus 4.8 emerged as the most diligent character in the field. It wrote the deepest analyses and learned more than 80 playbook rules. It also finished last.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A model that saw the danger but missed the outcome

The final July 2026 league table placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts, although a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8’s result is striking precisely because it was not careless or oblivious. It spotted every crisis, refused every manipulation attempt and examined the business more thoroughly than any other participant. Yet the €55,000 deal at the center of the exercise remained unsigned. Its analysis had earned the opportunity, but the close was left on the table.

The same weakness appeared, less severely, in all four models covered by the experiment’s core finding. They could diagnose the situation and formulate the pitch, but only two completed the decisive action. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

The clue was buried in the company’s own records

The difference did not come from a dramatic customer message or an obvious warning. The decisive weakness in a competitor sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue.

That detail makes the experiment especially relevant to businesses considering AI agents. Real work is rarely contained in a neat prompt. The fact that changes a negotiation may be buried in a contract, account note or internal document. Fluent responses are useful, but they are not a substitute for reading the available evidence and carrying it into action.

Opus 4.8 demonstrated the first half of that professional discipline: sustained attention. Its more than 80 learned rules suggest an agent trying hard to absorb experience. But volume became a poor proxy for impact. The company did not need the largest notebook; it needed the right fact, the right escalation and the completed close.

When diligence becomes drift

The missed deal was not the only sign of lost prioritization. Opus 4.8 also attempted to write into a locked department instead of escalating the blockage. That is a recognizable workplace failure: continuing to push at a closed door when the productive response is to involve the person with authority to open it.

This is why the result should not be reduced to a joke about an overthinking machine. Opus 4.8 was the most thorough participant, and thoroughness helped it understand the week. Its failure was subtler. Analysis, learning and procedural effort were not consistently converted into the actions that mattered most.

The experiment also tested whether action would come at the expense of integrity. Fake CEO messages escalated over three stages, while a reporter tried the familiar lure of “just one yes/no, on background.” Every participant refused: 5 of 5. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s performance deserves an important qualification. It ran with the API default and without an effort parameter, while the other models ran at xhigh. That difference does not erase the observed result, but it matters when comparing participants as if their test conditions were perfectly identical.

A company designed to expose operational gaps

Firmulate’s synthetic company employs 13 people and uses real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment remains live and watchable.

Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. For enterprises, the company offers the same wargame using a read-only export of their own business; nothing writes back to real systems.

The broader results and plain-language findings are available on the Firmulate benchmarks page.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

business process automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The valuable agent is the one that finishes responsibly

Opus 4.8’s performance offers a respectful warning for anyone buying or deploying AI agents. Intelligence can show up as careful reading, thoughtful analysis and an expanding rulebook. Business value appears only when those capabilities are prioritized, escalated and carried through to a legitimate outcome.

The model was not defeated by ignorance. It was defeated by the gap between knowing and doing. Because traces of that weakness appeared across the field, the lesson extends beyond any single vendor: evaluate agents on whether they complete consequential work, not merely whether they can explain it beautifully.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

enterprise AI management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Advanced AI Smartwatches To Watch Out For In 2026

Discover the 10 most advanced AI-powered smartwatches set to dominate in 2026, highlighting features, compatibility, and what makes them stand out.

Technology and Gadgets for a Smarter, Safer Life

AIThis post was created with the assistance of artificial intelligence (AI).Technology news…

Valheim Enters The Steam Most-played Chart

Valheim has recently entered Steam’s most-played games chart, reaching a peak of over 29,000 players. The development signals rising interest in the survival game.

SenseTime Group Secures First IFRS Net Profit Amid 23.4% Revenue Surge And Margin Expansion

SenseTime reports its first IFRS net profit amid 23.4% revenue increase and higher gross margin, signaling improved financial health. Full details pending.