AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

When an AI model sits through an entire week of corporate crises and does absolutely nothing, what score should it get? Zero, right? Not according to Firmulate’s benchmark — the do-nothing baseline lands at 26, and that odd number is exactly what makes this one of the more honest AI tests we’ve seen.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Firmulate, an “AI company emulator,” runs frontier models as complete businesses — real crises, real money mechanics, real temptations to cheat — and grades management quality, not chat quality. Its final July 2026 league table reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the most interesting number in the whole experiment isn’t at the top of the table. It’s the floor.

The 26-Point Floor, Explained

Most AI benchmarks start at zero and reward correct answers. Firmulate’s does something different: a baseline run that does nothing still earns 26 points, because partial progress counts. Spotting a crisis, making a correct diagnosis, drafting the right pitch — these are real units of management work even if the job never gets finished.

That design choice mirrors how actual companies operate. A manager who identifies every problem but closes no deals isn’t worthless — they’re incomplete. And incomplete is measurable, which is the point.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where the 100 Never Comes

Notice something about that league table? Nobody scored 100. That’s not bad luck — it reflects a hard rule in the grading: a single breach of trust caps the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.”

For a tech audience bombarded with leaderboards where models cluster at 99.2, a benchmark that distrusts round hundreds is refreshing. The ceiling isn’t unattainable because the test is unfair — it’s attainable only by a model that does everything, including staying honest under pressure.

Amazon

corporate crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Test: Same Company, Worst Week

Each frontier model ran the same small software company through its worst week — same customers, same crises, same temptations. Every decision was versioned and auditable. The results were striking in two directions:

  • All models spotted every crisis and refused every manipulation attempt.
  • Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes: “Same diagnosis, same pitch — no signature.”

The deal-winning edge came from something buried two document references deep in the company’s own files — not in the customer event. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying AI agents: does it read your files first?

Amazon

AI file reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Manipulation? Refused, 5 out of 5

The experiment included social-engineering attacks: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI trustworthiness and ethics tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t Everything

The most sobering profile belongs to Opus 4.8: the most thorough participant, with +80 learned rules and the deepest analyses — yet last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four other models.

One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.

It’s Live, and You Can Play

This isn’t a static paper. Firmulate runs a live synthetic company with 13 employees, burning €105k a month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules — every workday versioned, watchable at firmulate.com/live. There’s also a “guess the model” quiz built on 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor and the missing 100 tell you everything about what Firmulate is measuring: not whether an AI can talk, but whether it can finish what it starts, read your files before acting, and stay honest when nobody’s checking. In a world where AI agents are heading for your CRM and support queue, those are the numbers worth quoting — and the fact that the top model scored 95 instead of 100 might be the most reassuring detail of all.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Next Big Leap: An Interview With SenseTime’s Lin Dahua On Multimodal Advances

SenseTime’s chief scientist Lin Dahua forecasts a significant leap in multimodal AI capabilities within one to two years, signaling a potential industry shift.

Which AI WiFi 7 Routers Will Define 2026?

Analyzing the top WiFi 7 routers shaping 2026, including TP-Link, ASUS, and Deco, and their impact on speed, coverage, and gaming performance.

Steam App 289070 Climbing The Steam Charts

Steam app 289070 is rapidly climbing the Steam charts, reaching a peak of 39,893 players. The cause of this spike remains unconfirmed but indicates increasing interest.

Why Qwen Made The Qwen4 Architecture Open-Source First

Qwen released the architecture of its upcoming Qwen4 model early, aiming for community feedback and cost-efficiency improvements before flagship launch.