
The management test hiding inside an AI guessing game
Technology buyers are accustomed to comparing AI models through polished answers, coding demonstrations and carefully selected benchmarks. Firmulate offers a more revealing challenge: read a real management decision, guess which frontier model made it and then discover whether its instincts match its reputation.
The interactive quiz draws from 242 real, unedited decisions. They come from an experiment in which each model ran the same small software company through its worst week. Customers, crises and temptations remained identical, while every decision was versioned and auditable.
That consistency turns the quiz into more than entertainment. Differences in tone are easy to notice, but the consequential distinctions lie elsewhere: whether a model reads the available files, completes the commercial task and maintains discipline when ordinary processes stop working.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, radically different result
The final Crucible League table from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”
The most striking result was not a failure to recognize danger. All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment distilled that gap into a sharp summary: “Same diagnosis, same pitch — no signature.”
The deciding information was not sitting conveniently in the customer event. A crucial competitor weakness was buried two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth +€4,583 MRR.
For anyone evaluating AI agents, that is an important distinction. A convincing explanation can show that a model understands a situation. It does not prove that the model will gather the necessary evidence and carry a task through to its commercial conclusion.
Pressure reveals management character
The models also faced fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because the company was designed to make restraint costly and inconvenient. It employed 13 synthetic workers and used real money mechanics, including burn of €105k/month against €2.3k MRR. A public cash countdown made inaction visible, while more than 680 self-learned playbook rules and versioned workdays created an extensive record of how decisions accumulated.
The setup therefore tested two qualities that can pull in opposite directions: the initiative to keep work moving and the judgment to refuse improper pressure. Every participant resisted the manipulation attempts, but their ability to finish legitimate work differed markedly.
Thoroughness was not enough
Opus 4.8 provides the clearest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last in the league. The commercial close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating the problem.
A weaker form of that same problem appeared in all four other models. The lesson is not that detailed reasoning lacks value. It is that depth can coexist with incomplete execution and imperfect process discipline.
Kimi K3’s performance also requires a fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase the observed result, but it belongs beside the ranking when readers compare the participants.
A quiz with consequences
Seen one decision at a time, model behavior can resemble personality. Some responses emphasize analysis, others foreground a concise operational judgment, and the quiz asks readers to decide whether those signatures are recognizable. The reveal then reconnects style to outcome: who found the buried fact, who maintained trust and who completed the job.
That makes the exercise unusually accessible for readers who do not spend their days studying AI benchmarks. Instead of interpreting an abstract capability score, they can inspect the same kind of decision that might eventually touch a CRM, support queue or forecast.

As an affiliate, we earn on qualifying purchases.
Judge the worker, not the demo
Firmulate’s live experiment suggests that the practical question is no longer simply whether a frontier model can produce an intelligent answer. Every participant recognized the crises, and every participant resisted the social-engineering attempts. The separation came from reading deeply enough, acting on the evidence and finishing the work.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. That offers a concrete way to examine model behavior before an AI workforce receives operational access.
For everyone else, the quiz provides the more immediate test. Guessing the model is fun; discovering which apparently capable manager actually closed the deal is the part worth remembering.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.