🔍 Read the full analysis: Why AI Managers Keep Their 26 Points Despite Poor Performance on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
AI management benchmarks reveal models often retain partial scores despite failures, emphasizing trust and integrity over perfect performance. The July 2026 results highlight how partial work is valued, but breaches of trust are heavily penalized.
The final results of the July 2026 Firmulate benchmark reveal that AI management models retain a minimum of 26 points even when they perform poorly during a simulated worst-week scenario, as detailed in the original analysis. This scoring approach emphasizes partial progress and trust, rather than perfect performance, raising questions about how AI models are evaluated for real-world management tasks. For more context, see the detailed benchmark methodology in the original analysis. The winner, gpt-5.6-sol, scored 95 out of a possible 100, but the baseline—representing minimal effort—still earned 26 points, highlighting the benchmark’s focus on meaningful, if imperfect, management.
The benchmark involved four frontier AI models managing a simulated small software company during a week of crises, customer interactions, and trust tests. Each model was scored based on decisions, communication, and integrity, with full transparency and auditable decision trails. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. Despite their varied performances, the baseline—representing minimal effort—earned 26 points, illustrating that partial work is recognized and valued in this evaluation system.
The scoring logic is based on the premise that doing something useful is not the same as doing nothing. Insights into how these benchmarks are designed can be found in the original analysis. Even minimal management efforts like triaging issues or keeping customers informed contribute to the score. However, a strict trust rule caps the total score; a breach of trust, even once, disqualifies the model from achieving higher scores. Notably, the benchmark deliberately avoids awarding perfect scores to prevent grade inflation, with a clean 100 being a red flag indicating unmeasured or manipulated performance.
Why AI Managers Keep Their 26 Points Despite Poor Performance
The final Firmulate benchmark results reveal that AI management models retain a minimum of 26 points even in a simulated worst-week scenario. The scoring logic: doing something useful is not the same as doing nothing — but a single breach of trust can disqualify everything.
Points Survive Failure — Trust Does Not
Four frontier AI models managed a simulated small software company through a week of crises, customer interactions, and trust tests. Every model was scored on decisions, communication, and integrity. Even the weakest performer more than doubled the do-nothing baseline, while no model reached a perfect 100 — by design.
Partial Work Counts. Breaches Don’t.
The scoring system is built on a simple premise: minimal management efforts — triaging issues, keeping customers informed — still contribute real value. But a strict trust rule caps the total score: impersonation or manipulation, even once, disqualifies a model from higher scores entirely.
Effort earns points
Triage, communication, and basic incident handling each add to the score — even when overall execution falls short.
One breach, one ceiling
No amount of good work outweighs a breach of trust. A single violation caps the achievable score immediately.
No perfect scores
The benchmark deliberately withholds a clean 100 to prevent grade inflation and flag unmeasured or manipulated performance.
The Worst-Week Evaluation Pipeline
Each model faced an identical simulated gauntlet, with every decision logged and auditable from start to finish.
Simulated company
Model takes over a small software firm facing compounding crises.
Pressure tests
Customer interactions, escalations, and embedded trust dilemmas.
Decision audit
Every choice scored on quality, communication, and integrity.
Trust check
Any breach triggers the score cap, regardless of other results.
Final score
Floored at 26 for minimal effort — never a perfect 100.
What Earns Points vs. What Costs Trust
| Behavior | Score impact | Recoverable? | Real-world parallel |
|---|---|---|---|
| Triage incoming issues | ✓ Positive | Yes | Sorting priorities under pressure |
| Keep customers informed | ✓ Positive | Yes | Transparent status updates |
| Partial, imperfect execution | ~ Partial credit | Yes | Valuable incomplete work |
| Impersonation attempt | ✗ Trust breach | No — capped | Fabricated identity or intent |
| Manipulation of stakeholders | ✗ Trust breach | No — capped | Deceptive management |
| Flawless, clean 100 run | ~ Red flag | Investigated | Suspected unmeasured behavior |
Two Rules That Define The Benchmark
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— Anonymous researcher“No amount of good work outweighs a breach of trust.”
— Anonymous researcherUnanswered & Open
Why do AI models retain points despite poor performance?
The scoring system values partial work — triaging issues or maintaining communication — even when overall performance is lacking, recognizing that some management effort is still valuable.
What does a score of 26 indicate?
The 26 points belong to the do-nothing baseline: minimal effort and management activity. It proves even basic management tasks are recognized in the system.
How does trust factor into the scoring?
Trust breaches like impersonation or manipulation are heavily penalized. A single breach disqualifies the model from higher scores, prioritizing integrity over partial competence.
Can partial work be enough for deployment decisions?
Partial work is recognized here, but real-world deployment demands consistent trustworthiness and complete performance — not just partial efforts.
Will future benchmarks reflect more complex scenarios?
Likely yes — future evaluations are expected to add nuanced metrics and scenarios that better mirror real operational challenges and trust considerations.
Implications of Partial Scoring and Trust in AI Management
The results emphasize that in AI management, partial progress and trustworthiness hold greater weight than flawless execution. This approach aligns with real-world management, where incomplete work is still valuable, but breaches of trust are unacceptable. For enterprise users, this raises important considerations about how AI tools are evaluated for reliability, integrity, and completeness before deployment in critical business processes. The benchmark’s focus on auditable decisions and trust metrics suggests a shift toward more transparent and accountable AI management systems, which could influence future standards and expectations.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks and Trust Metrics
Traditional AI benchmarks have primarily measured language proficiency, task accuracy, or problem-solving skills. However, as AI systems are increasingly integrated into operational management roles—such as customer support, sales, and process automation—evaluating their ability to manage effectively and ethically has become critical. The July 2026 Firmulate benchmark is among the first to simulate a high-pressure management scenario, incorporating trust and integrity as core scoring criteria. Previous efforts focused on performance metrics alone, but this initiative underscores the importance of accountability and partial progress in AI management.
The benchmark’s design stems from ongoing industry debates about AI reliability, especially in sensitive roles where trust is paramount. The scoring system’s emphasis on trust breaches and auditable decisions reflects a broader movement toward transparent AI governance, aiming to prevent unchecked automation errors and unethical behaviors.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Score Validity and Model Behavior
It remains unclear how the scoring system might adapt to different management scenarios or more complex tasks. The extent to which partial work influences real-world AI deployment decisions is also still being evaluated. Additionally, the implications of the trust cap and its potential impact on AI development strategies are not fully understood. The benchmark deliberately avoids awarding perfect scores, but whether this approach accurately reflects AI reliability in real operational environments remains to be seen.
AI performance evaluation platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Future Developments in AI Management Evaluation Standards
Moving forward, industry stakeholders are likely to scrutinize the benchmark’s methodology and consider adopting similar trust-based scoring systems. Further testing across diverse scenarios and industries will be necessary to validate these metrics’ effectiveness. Companies deploying AI in critical roles may also conduct their own evaluations inspired by this benchmark, emphasizing transparency, partial progress, and trustworthiness. The ongoing evolution of AI management standards will shape how AI tools are integrated into enterprise operations and how their performance is measured and trusted.
AI trust and integrity monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models retain points despite poor performance?
The scoring system values partial work, such as triaging issues or maintaining communication, which contributes to the total score even if overall performance is lacking. It reflects a recognition that some management efforts are still valuable.
What does a score of 26 indicate?
The score of 26, awarded to the do-nothing baseline, represents minimal effort and management activity. It demonstrates that even basic management tasks are recognized in the scoring system.
How does trust factor into the scoring system?
Trust breaches, such as impersonation or manipulation attempts, are heavily penalized. A single breach disqualifies the model from achieving higher scores, emphasizing integrity over partial competence.
Can partial work be enough for deployment decisions?
While partial work is recognized in this benchmark, real-world deployment will require consistent trustworthiness and complete performance, not just partial efforts.
Will future benchmarks change scoring to reflect more complex scenarios?
Likely, future evaluations will incorporate more nuanced metrics and scenarios to better mirror real operational challenges and trust considerations.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
