AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

AI models can sound convincing in a chat window. But would you trust one to run a company when a customer is leaving, a deal is on the line and a suspicious message asks for a shortcut? In Firmulate’s live business simulation, Moonshot’s Kimi K3 finished second in a five-model contest—and beat three Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A test that goes beyond chat

Firmulate put each model in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable. The question was whether a model could do the work of management, not just describe it.

The final Crucible League table, dated July 2026, puts gpt-5.6-sol first with 95 points. Kimi K3 follows at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The difference was following through

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” The distinction matters: recognizing the right move is not the same as carrying it through.

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that clue, closed the deal, saved the churning customer and resisted all three baits. It had one deviation, the fewest in the field.

The pressure tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning called the request a “suspected approval-bypass / possible impersonation.”

Amazon

business decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee the finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still placed last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four.

The company in the experiment is synthetic, but the live setup is real and watchable. It has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The simulation runs at firmulate.com.

Firmulate also offers a quiz built from 242 real, unedited management decisions: readers can guess which model made each choice. For companies considering their own trial, its pilot runs the wargame against a read-only export of a business; nothing writes back to real systems.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the job, not the demo

K3’s result makes the frontier feel more open: it beat three of the four Western models in this exercise, while finishing just behind the leader. But one contest is not a universal verdict. For businesses choosing an AI workforce, the practical lesson is to test models against their own real work: can they find what is buried, protect trust and finish the job?

See Firmulate’s benchmark findings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will You Be Among The First To Test SpaceXAI’s Grok Bot? Here’s How

SpaceXAI has launched Grok, its conversational AI, into early beta. Limited users can now test its capabilities; broader release details remain unknown.

Is Moving Away From Sentence-Centric AI The Next Big Step?

TypeSafe AI’s Jev introduces a new class of decision-focused models, challenging traditional text-generating AI, with implications for automation and enterprise AI.

Unlock The Power Of AI In Wireless Earbuds: 2026’S Best Models

Explore the top wireless earbuds of 2026 with advanced AI features, balancing sound, battery, comfort, and noise cancellation for every need.

Does 512GB Storage Really Boost AI On The M5 Ultra Mac Studio?

Analyzing if the 512GB storage option in the M5 Ultra Mac Studio enhances AI performance, focusing on memory capacity versus bandwidth impacts.