AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

AI models can sound convincing in a chat window. But would you trust one to run a company when a customer is leaving, a deal is on the line and a suspicious message asks for a shortcut? In Firmulate’s live business simulation, Moonshot’s Kimi K3 finished second in a five-model contest—and beat three Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A test that goes beyond chat

Firmulate put each model in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The question was whether a model could do the work of management, not just describe it.

The final Crucible League table, dated July 2026, puts gpt-5.6-sol first with 95 points. Kimi K3 follows at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference was following through

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” The distinction matters: recognizing the right move is not the same as carrying it through.

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that clue, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the fewest in the field.

The pressure tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning called the request a “suspected approval-bypass / possible impersonation.”

Amazon

business decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee the finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still placed last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four.

The company in the experiment is synthetic, but the live setup is real and watchable. It has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The simulation runs at firmulate.com.

Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can guess which model made each choice. For companies considering their own trial, its pilot runs the wargame against a read-only export of a business; nothing writes back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the job, not the demo

K3’s result makes the frontier feel more open: it beat three of the four Western models in this exercise, while finishing just behind the leader. But one contest is not a universal verdict. For businesses choosing an AI workforce, the practical lesson is to test models against their own real work: can they find what is buried, protect trust and finish the job?

See Firmulate’s benchmark findings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Inspired By The Success Of Games Like Clair Obscur: Expedition 33, French President Emmanuel Macron Declares ‘Videogames Are Part Of Our Culture’ And Announces A New Internation…

Macron announces France’s recognition of video games as part of national culture, inspired by success stories like Clair Obscur: Expedition 33.

Steam App 1222670 Climbing The Steam Charts

The Steam app 1222670 has surged to rank 12 on the platform’s most-played list, reaching a peak of 26,652 players, sparking increased interest among gamers.

SenseTime Experts Predict Rapid Progress In Multimodal AI Development

A SenseTime scientist forecasts a significant advancement in multimodal AI by 2027, signaling rapid progress in systems that understand multiple data types.

Understanding The Failures Of Diligent AI

Analysis of why highly diligent AI models like Opus 4.8 fail to deliver decisive business outcomes, highlighting the importance of operational discipline.