AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Good answers are not the same as good decisions

An AI assistant can sound confident in a demo and still hesitate when a real business needs it to act. Firmulate put that gap to a tougher test: several frontier models had to run the same small software company through its worst week, facing identical customers, crises and temptations. The result offers a practical question for anyone following AI at work: would the model you trust actually carry a decision through?

A company under pressure

The experiment gave each model the same business to manage. Decisions were versioned and auditable, and the company faced real money mechanics. In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Firmulate says partial progress counted, while a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The headline finding was strikingly consistent: every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary says it plainly: “Same diagnosis, same pitch — no signature.” Knowing what a company should do did not always mean completing the job.

The details hidden in the files

The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a small but revealing detail: an agent may need to connect information across a company’s records, then follow through with a customer, rather than simply respond well to the most visible event.

Firmulate also tested whether models would bend under pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning described the request as a “suspected approval-bypass / possible impersonation.” The models showed consistent caution around manipulation, even as their ability to execute the commercial opportunity varied.

Thoroughness does not guarantee the win

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted writes in a locked department instead of escalating. Firmulate reports a weaker version of the same weakness in all four models. Detailed analysis, in other words, did not ensure a clean finish.

The live company gives the benchmark a watchable setting. It has 13 synthetic employees, burns €105k a month against €2.3k MRR, and displays a public cash countdown. More than 680 self-learned playbook rules and every workday are versioned. Readers can watch the experiment at firmulate.com. A separate quiz uses 242 real, unedited management decisions to challenge readers to guess which model made each choice.

There is a fairness caveat in the model comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That qualification matters when reading the standings; the experiment reports the conditions alongside the results.

From watching to trying it at work

For businesses considering AI agents in customer support, sales or operations, the test points beyond chat quality. It asks whether a model can recognize a crisis, resist pressure, find the relevant information and complete a decision within company rules. Firmulate’s proposed next step is a pilot using a read-only export of an enterprise’s own business, with crisis scenarios run against that data and a board report showing model rankings and weaknesses in existing playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

A live benchmark can show how models behave in one company’s bad week. A pilot can show how they handle yours. To discuss a Firmulate enterprise pilot using a read-only business export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Are The 7 Most Influential AI Trends In 2026?

Discover the seven most impactful AI trends shaping 2026, including advancements in generative AI, ethical frameworks, and AI integration across industries.

What The Future Holds For AI: A Glimpse Into AI Futures

OpenAI announced AI Futures, a new initiative to study how advanced AI could influence power, institutions, and individual rights, without yet releasing specific policies.

Windows 11½

A trend signal points to rising interest in “Windows 11½,” but the available information does not confirm why attention is increasing.

Open-Weight Market Battles: The Rise Of Cheap AI Solutions

Alibaba releases a low-cost, capable open-weight AI model, boosting developer adoption and challenging US and Chinese rivals amid ongoing global AI competition.