AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Coding skill is not management judgment

Technology buyers have learned to scan coding leaderboards and chat arenas for signs of the smartest model. Those tests can reveal answer quality, but they leave a more consequential question unresolved: what happens when an AI agent must choose among competing emergencies, operate under capacity pressure and remain candid when the news is bad?

That gap matters once an agent moves beyond drafting text and begins touching a support queue, CRM or forecast. A persuasive response is not the same as a resolved customer crisis. Correct analysis is not a signed contract. Refusing an obviously improper request once is not the same as preserving trust across days of pressure.

Firmulate is turning that distinction into a public experiment. Its premise is that businesses should measure management quality, not chat quality.

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The company becomes the benchmark

In the Crucible League, each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained constant; only the model changed. Every decision was versioned and auditable.

The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the exercise imposed a hard principle on trust: “no amount of good work outweighs a breach of trust.” A single breach therefore capped the total.

Those results look like a familiar ranking, yet the underlying stories are more useful than the order. Every model spotted every crisis and rejected every manipulation attempt. Only two signed the €55,000 deal that their own analysis had earned. The finding can be reduced to a painfully recognizable management failure: “Same diagnosis, same pitch — no signature.”

This is where conventional evaluations lose the plot. An agent can understand the situation, prepare the right material and still fail to complete the commercially decisive action. In a chat window, an excellent proposed pitch may look like success. Inside a company, the unsigned deal remains unsigned.

Reading the company before reacting

The deal also exposed the value of organizational memory. The decisive weakness in a competitor was not present in the customer event. It sat two document references deep in the company’s own files. Models that found and used that fact won the contract at full price, worth +€4,583 MRR.

That is a management lesson disguised as an information-retrieval task. Real work rarely arrives as a self-contained prompt. The most important context may be buried in account history, an earlier decision or a document connected indirectly to the event. A model that reacts fluently to the latest message can still miss the fact that changes the outcome.

Trust held up better than execution

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

This finding complicates the easy narrative that safety and usefulness sit at opposite ends of a trade-off. In this field, the models consistently resisted manipulation. The larger separation came from follow-through, research depth and operational discipline.

Opus 4.8 makes that point especially clearly. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The commercial close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other models, though less strongly.

Thoroughness, then, is not a substitute for completion. An organization does not receive the value of an analysis merely because the analysis exists. The agent must recognize the next authorized action, take it and handle obstacles through the proper route.

There is also an important fairness note around Kimi K3’s result: it ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs beside the score rather than buried beneath it.

A curriculum of consequences

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real, public and watchable, rather than a fictional case study reconstructed after the fact.

The scenario names point toward a new evaluation curriculum: churn wave, price increase, downround and PR crisis. Each forces choices whose consequences continue beyond a single response. They test whether an agent can prioritize, retrieve context, close loops and tell the board the truth when optimism would be easier.

Readers can inspect the full benchmark results, while 242 real, unedited management decisions also power a “guess the model” quiz. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI trust and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers need evidence of finished work

The next useful AI leaderboard will not discard coding tests or chat comparisons. It will put them in context. Answer quality remains valuable, but it is only the opening move when an agent is expected to operate a business process.

Before hiring an AI workforce, buyers should ask what happens after the polished response: Did the agent read the relevant files? Did it distinguish urgency from importance? Did it complete the authorized action? Did it resist pressure without becoming inert? Did it surface bad news honestly?

Firmulate’s results suggest that frontier models can share a diagnosis while producing sharply different business outcomes. That is the measurement gap technology leaders now need to close. The decisive benchmark is no longer whether an AI sounds like a capable manager. It is whether the company is better managed when the conversation ends.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI organizational memory retrieval systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The August 2026 AI Update: Everything You Need To Stay Informed

Google’s August 2026 AI announcements include Gemini 3.7 Flash, Pixel 11 series, and Gemini app surpassing 1 billion users. Here’s what you need to know.

SenseTime’s Swing To First-Half Profit Could Be A Game Changer For SenseTime Group (SEHK:20) – Simplywall.st

SenseTime Group reportedly swung to a first-half profit, a move that could alter investor perceptions, but key details remain unconfirmed and unclear.

Technology and Gadgets for a Smarter, Safer Life

AIThis post was created with the assistance of artificial intelligence (AI).Technology news…

Grand Theft Auto 6 Leaks Response

Rockstar Games has issued a statement following the recent massive leak of Grand Theft Auto 6 gameplay footage and assets, confirming they are investigating the breach.