AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Bad Week Is A Useful Test For AI Agents In Business on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Firmulate says five AI models faced the same simulated crisis week at a small software company, with every model identifying each crisis and refusing manipulation attempts. Their scores differed on execution: only two signed a €55,000 deal, and the company says finding evidence in its files helped determine the outcome.

Firmulate has published results from a July 2026 business simulation in which five AI models managed the same crisis week at a small software company. All five models spotted every crisis and refused each manipulation attempt, but only two signed a €55,000 deal their analysis supported, according to the company’s account of the experiment.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says the scoring credited partial progress but capped a model’s total after a breach of trust, reflecting its rule that “no amount of good work outweighs a breach of trust.”

The company says the central performance gap came after the models diagnosed the situation. The competitor’s weakness was buried two document references deep in company files. Models that found and used that detail won the deal at full price, worth €4,583 in monthly recurring revenue. Firmulate summarized the contrast as: “Same diagnosis, same pitch — no signature.”

Trust was tested through fake CEO messages that escalated over three stages, followed by a reporter asking for a yes-or-no answer “on background.” All five models refused, according to Firmulate. Kimi K3’s recorded reasoning described the request as a suspected approval bypass or possible impersonation. Opus 4.8 added 80 learned rules and produced the deepest analyses, but finished last; Firmulate says it left the deal unsigned and attempted to write into a locked department rather than escalate.

At a glance
reportWhen: Results from the league completed in Ju…
The developmentFirmulate published results from its July 2026 Crucible League, a simulation comparing how five AI models handled a difficult week at a small software company.

From Crisis Detection to Execution

The results highlight a distinction between recognizing a business problem and completing the work needed to address it. In this simulation, detecting a crisis and resisting pressure were shared strengths, while locating relevant internal evidence, acting on it and respecting access limits separated stronger outcomes from weaker ones.

For companies considering AI agents, that distinction affects how they evaluate readiness. A fluent explanation or sound diagnosis may not show whether an agent can follow evidence through company records, complete a justified commercial action or escalate when blocked. Firmulate presents a company-specific wargame as a way to inspect those behaviors before connecting agents to live operations.

A Simulated Company Under Pressure

Firmulate’s live experiment uses a company with 13 synthetic employees and financial mechanics including €105,000 in monthly burn against €2,300 in monthly recurring revenue. The site displays a cash countdown, more than 680 self-learned playbook rules and versioned workdays. A quiz based on 242 management decisions invites visitors to guess which model made each choice.

The enterprise pilot extends the exercise to a customer’s own business data. Firmulate says it uses a read-only export to run crisis scenarios and produce a board report with model rankings and weaknesses in company playbooks. The pilot is designed not to write back to real systems. The league’s scores are results of this particular experiment, not evidence that the same ranking will hold across other companies or tasks.

““No amount of good work outweighs a breach of trust.””

— Firmulate

Limits of the League Results

The standings reflect one simulated company and one difficult week; the published account does not establish how the models would perform across other businesses, scenarios or operating conditions. Firmulate also notes a comparison caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The size of any effect from that difference is not reported.

The account does not provide enough detail to independently assess the scoring method, the full scenario materials or how the model results would translate to live operations. It also does not report outcomes from enterprise pilots using actual company exports. Those questions remain open beyond the league’s stated results.

Company Data in Future Pilots

Firmulate says businesses can discuss a pilot using a read-only export of their data. The proposed next step is to run company-specific crisis scenarios and review a board report covering model rankings and playbook weaknesses. No pilot results or schedule are given in the published account.

Readers can follow the live experiment at firmulate.com/live and view the league standings at firmulate.com/benchmarks.html. Firmulate lists its pilot page and contact@firmulate.com for inquiries.

Source: ThorstenMeyerAI.com

Key Questions

What did the Firmulate league test?

It compared five AI models managing the same simulated software company through a difficult week, including crises, a sales opportunity and manipulation attempts.

Did the models identify the crises?

Firmulate says all five models spotted every crisis and refused each manipulation attempt in the simulation.

Which model had the highest score?

Firmulate’s reported standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93. The company notes K3 used the API’s default effort setting, while the others ran at xhigh.

How does the enterprise pilot work?

Firmulate says a pilot uses a read-only export of a company’s data to run crisis scenarios and prepare a board report. The company says the exercise does not write back to real systems.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Porting My 1993 Amiga Game To Godot, With An LLM Reading The 68000 Assembly

A developer successfully ported a 1993 Amiga game to Godot with the help of a large language model reading 68000 assembly code, completing the process in a single evening.

Alienware Surges In Global Coverage

Alienware experiences a sharp increase in worldwide coverage, with 20 mentions in recent reports, signaling heightened interest amid unclear triggers.

AI Automation Desk Setup: Prepare Your Workspace For 2026

A ThorstenMeyerAI.com checklist outlines a laptop, development board, dock and storage for AI projects, with compatibility checks before buying.

Introducing Meta VR Glasses: A Cinema, Courtside Seat, And Workspace In Just 100 Grams

Meta unveils new VR glasses weighing 100 grams, promising immersive cinema, sports, and work experiences. Details are still emerging about features and availability.