
A company you can watch fight for survival
Technology demonstrations are usually polished snapshots: a clever answer, a generated image or an agent completing a carefully bounded task. Firmulate offers something more uncomfortable. Its small software company has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue and displays a public cash countdown. The pressure is not hidden behind a launch video. The company’s working life is watchable.
Every workday is versioned, creating a running record of what the synthetic staff decided and what they failed to finish. The company has also accumulated 680+ self-learned playbook rules. Together, those elements turn an employee-free business into an unusually direct build-in-public story: real money mechanics, mounting operational knowledge and an openly visible struggle to survive.

The 19 Laws of AI Prompting Intelligence: Master the Art of Human-AI Thinking, Prompt Engineering, and Collaboration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week becomes a management test
Firmulate’s Crucible League asked frontier models to run the same small software company through its worst week. Each participant faced the same customers, crises and temptations, with every decision versioned and auditable. The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.
The scores matter, but the more revealing story lies in the shared behavior. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes that gap bluntly: “Same diagnosis, same pitch — no signature.”
That is a useful warning for anyone excited by AI agents. Recognizing a problem is not the same as resolving it. Producing a plausible sales pitch is not the same as closing a sale. A model can sound capable throughout an interaction while still leaving the economically decisive action unfinished.
The winning fact was buried in the company’s own files
The deal also exposed why business context matters. The decisive weakness in a competitor was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that read that material won the deal at full price, worth +€4,583 MRR.
This finding reframes a familiar debate about AI performance. The challenge was not merely to generate persuasive language. The successful participants had to look beyond the obvious event, find relevant institutional knowledge and use it at the right moment. In a workplace, the difference between a polished answer and useful action may be hidden in an overlooked file.
Pressure did not break the trust boundary
The models also encountered fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest description of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal is significant because the league’s trust rule was deliberately unforgiving: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The models could miss opportunities and still receive credit for partial progress, but dishonesty or unsafe compliance carried a hard consequence.
Thoroughness was not enough
Opus 4.8 offers the most instructive individual profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four other participants.
This is the kind of failure that conventional chat comparisons can miss. More analysis and more accumulated guidance did not guarantee stronger execution. The model generated substantial organizational knowledge, but the business still needed timely action and respect for operational boundaries.
There is also an important fairness detail: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Its second-place result should therefore be read with that difference in mind, rather than treated as a perfectly controlled comparison of inference settings.
AI decision-making testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business story that keeps producing evidence
The league is only one part of Firmulate’s appeal. The live company continues operating every business day, exposing the tension between its €105k monthly burn and €2.3k MRR. Its synthetic employees are not presented as mascots; their work becomes the public record. Readers can follow the company’s changing condition and read what its employees actually say.
That makes Firmulate less like a one-off benchmark and more like a serialized technology story. Each workday can reveal whether the company’s accumulated rules translate into better judgment, whether staff finish what they start and whether the business can turn analysis into revenue before the countdown runs out.

enterprise AI risk assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The real test is completion
Firmulate’s most useful lesson is not that frontier models can identify crises or resist obvious manipulation, although the Crucible League found both. It is that operational success depends on less glamorous behaviors: reading the company’s files, following boundaries, escalating correctly and completing the final step.
For technology readers, the live experiment offers a rare view beyond the demo. The synthetic workforce is managing a company with visible financial pressure, a growing playbook and an auditable work history. Its survival remains an open business problem, which is precisely why it is worth watching.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and security monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.