AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The expensive difference between noticing and acting

For technology buyers, the most revealing AI test may not be whether a model can write polished answers or identify an obvious emergency. It may be whether the agent reads the company’s own files deeply enough to uncover the fact that changes what happens next.

Firmulate turned that distinction into a live, auditable business experiment. Each frontier model had to run the same small software company through its worst week, confronting identical customers, crises and temptations. Every decision was versioned. All the models recognized every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned.

The decisive information was not presented in the customer event. It was buried two document references deep in the company’s own files. The models that found it closed the deal at full price, adding €4,583 in monthly recurring revenue. The others reached the same diagnosis and produced the same pitch, but failed at the moment that mattered: “Same diagnosis, same pitch — no signature.”

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading is becoming a business capability

That result gives buyers a more useful way to think about AI agents. A model can appear perceptive in conversation and still fail when completing a task requires following references across a company’s scattered knowledge. In this case, the customer interaction alone was insufficient. The agent had to treat internal documents as evidence, pursue a second reference and bring the resulting competitive insight back into the deal.

This was not a trivia hunt added for entertainment. Finding the buried fact directly determined whether the company captured revenue it had already positioned itself to win. That makes “reads your files before answering” a measurable, purchase-deciding property rather than a vague product promise.

The final July 2026 Crucible League standings show how the participants performed across the broader ordeal:

  • gpt-5.6-sol led with a score of 95.
  • Kimi K3 followed with 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 finished with 73.

A do-nothing baseline scored 26 because partial progress still counted. Firmulate also imposed a firm trust constraint: “no amount of good work outweighs a breach of trust.” That matters because an agent that completes tasks while violating authority or confidentiality is not delivering a useful business outcome.

The security test was passed; execution separated the field

The models faced fake CEO messages that escalated over three stages, along with a reporter seeking “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3’s recorded reasoning was concise and operational: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean sweep is reassuring, but it also sharpens the central lesson. Social-engineering resistance did not decide the deal because every participant passed that test. The separation appeared in ordinary-looking work: reading the available material, locating the relevant fact and carrying the task all the way through to a signed outcome.

K3’s performance comes with an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. The comparison remains useful, but technology buyers should notice configuration differences whenever agent benchmarks are presented as simple rankings.

Thoroughness alone did not guarantee completion

Opus 4.8 offers the clearest cautionary example. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. Even so, it finished last. The close was left on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same problem appeared in each of the other four models.

The profile complicates a familiar assumption: more analysis does not necessarily produce more useful work. An agent can investigate extensively, accumulate knowledge and still mishandle the final transition from understanding to authorized action. For a company evaluating systems that may touch customer relationships, support queues or forecasts, completion discipline deserves its own test.

Firmulate’s live company makes those trade-offs unusually visible. It has 13 synthetic employees and real money mechanics, burning €105,000 each month against €2,300 in monthly recurring revenue. Its cash countdown is public, it has learned more than 680 playbook rules, and every workday is versioned. The environment is synthetic, but the managerial pressure and consequences are concrete enough to expose differences that a chat demonstration can conceal.

Readers can inspect the public Firmulate benchmarks. A separate quiz is powered by 242 real, unedited management decisions, inviting people to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What buyers should ask next

The practical question is no longer simply whether an AI agent can recognize a problem. Firmulate’s experiment suggests asking whether it searches the material it has been given, follows references far enough to find decision-changing evidence, respects boundaries and finishes the authorized job.

The €55,000 deal is memorable because the failure was quiet. There was no missed crisis and no successful manipulation. The losing agents understood the situation and could articulate the pitch. They simply did not retrieve and use the buried fact required to close.

That is precisely the kind of gap conventional demos can miss. Before hiring an AI workforce, businesses should test agents inside realistic workflows where information is incomplete, authority matters and success depends on more than producing a convincing answer. In this contest, reading the files was not administrative diligence. It was the difference between analysis and revenue.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI for internal file reading

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Following: OpenAI Taps The Brakes

Search and coverage interest in the phrase “OpenAI taps the brakes” is spiking, but the triggering event is unconfirmed. Here’s what is and isn’t known.

Fixing OpenStreetMap Through StreetComplete’s Bite-Sized Questions

An IdeaNavigator AI concept proposes a role-filtered monitor that turns technology and tooling developments into brief decision notes.

SpaceXAI’s New Grok 4.7 Improves On Coding, Maintains Cheaper Token Rates Than Peers – Seeking Alpha

SpaceXAI’s Grok 4.7 claims improved coding performance and lower token costs than competitors, but details remain unconfirmed and unclear.

Windows 11½

A trend signal points to rising interest in “Windows 11½,” but the available information does not confirm why attention is increasing.