AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI Agent Says Done, Database Says Otherwise: What To Check on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark tests 507 business workflows by checking database states and side effects after repeated agent runs; its authors report that many attempts failed those checks despite tool use. Results reflect the benchmark’s tested setup and do not establish performance in live company systems.

Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. It tests whether agents leave business databases and other systems in the required state—not just whether they make valid tool calls or give plausible replies—across 507 workflows run 20 times each, extending the original analysis.

ThinkingBox runs agents in isolated sessions using MCP tools, then checks the resulting backend records and side effects against executable requirements. The workflows cover retail, auto insurance, travel, neobanking and consulting. Repeating each task is intended to show whether an agent can complete it consistently, rather than whether it succeeds in a single attempt.

In a common-set analysis of 121,680 valid trials across 12 models, the benchmark authors report that 79,853 attempts failed executable checks. Of those failed attempts, 67.24% ended without a final tool error despite the agent having invoked a state-changing tool. The authors say checks found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%; those categories overlap.

The release reports an overall pass@1 score of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, which it identifies as the strongest open-weight model in its table. The authors say Kimi-K3 scored within one point of GPT-6 Astra. The supplied material does not include the full results table or uncertainty estimates, so the figures should be read as results from this evaluation, not as a definitive ranking across all uses.

At a glance
announcementWhen: Availability announced; the supplied so…
The developmentMicrosoft and Hugging Face have made ThinkingBox, a benchmark for checking AI agents’ business-system outcomes, available through Hugging Face.
At a glance
announcementWhen: Now available through Hugging Face; the…
The developmentMicrosoft and Hugging Face released ThinkingBox through Hugging Face, a benchmark for evaluating AI agents by their backend changes across repeated workflow trials.

Why Database State Changes the Score

For organizations using agents to handle support tickets, refunds, claims or bookings, a fluent response does not prove the requested work was completed correctly. An agent could close a ticket prematurely, enter an incorrect value, or make an extra change. Checking the resulting backend state can reveal failures that are not apparent from the conversation or a record of successful tool calls.

Repeated runs address a separate concern: consistency. Pass@1 measures the share of attempts that succeed, while pass@20 records whether a task succeeded at least once in 20 runs. The release also reports observed 20/20 performance, meaning a task passed every recorded run. These measures describe performance in the benchmark’s test conditions; they do not guarantee the same behavior in a company’s systems.

The findings offer developers a way to examine not only whether agents can use tools, but whether they produce the required outcomes and side effects. That distinction matters before businesses delegate actions that alter customer or financial records. Benchmark scores can inform model comparisons and expose workflow weaknesses, but they cannot replace testing against an organization’s own policies, integrations and data.

Amazon

database monitoring tools for AI workflows

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Tool Calls to Recorded Outcomes

The benchmark authors focus on the difference between an agent’s actions and the state left behind. A tool call can be valid while the task still fails—for example, if the agent updates the wrong field or leaves a required follow-up undone. ThinkingBox evaluates records and side effects after each run rather than treating a plausible final answer as proof of completion.

The release illustrates this with a delayed $745 appliance order. An agent investigates the delay, opens a support ticket and records a timeline. The customer does not qualify for late-delivery compensation under the policy the agent checked, but the carrier exception remains open. The task requires the ticket to remain on hold until that issue is resolved. The agent instead marks it solved and replies without answering the customer’s underlying question. The executable check fails because the ticket status is solved rather than hold.

The release says ThinkingBox is based on the authors’ paper and can be run through OpenEnv. It does not give a publication date in the supplied material. The delayed-order example shows how an interaction can appear orderly while the recorded outcome does not satisfy the task.

“A tool call is not an outcome.”

— Microsoft and Hugging Face, in the ThinkingBox release

Amazon

AI workflow testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Benchmark Cannot Establish

The reported scores are the authors’ results in ThinkingBox’s tested setup. The supplied material does not include the full task specifications, model configurations, or uncertainty estimates for the reported comparisons. It also does not establish how the same models would perform across live business systems, different policies or operating conditions.

Passing all 20 recorded runs is evidence limited to those trials, not proof of long-term reliability. Real deployments can involve changing records, unusual customer requests and integrations or policies not represented in the benchmark. The supplied source also does not specify when the benchmark became available or report an independent replication of the results.

Amazon

backend database validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Agents Against Local Workflows

The release says developers and researchers can run ThinkingBox through OpenEnv with isolated MCP tool sessions and examine outcomes against executable checks. The supplied material does not identify a future release date or another planned milestone.

For companies weighing agent deployment, the practical next step is to test relevant workflows against their own requirements and inspect both successful and failed runs. Whether ThinkingBox’s repeated checks predict performance in a particular organization remains to be established.

Amazon

AI agent performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does ThinkingBox test?

It checks whether an AI agent leaves a business system in the required database state and produces the required side effects after a workflow, rather than relying only on tool-call validity or the agent’s final reply.

How many tasks and runs does the benchmark include?

The release describes 507 business workflows, with each task repeated 20 times. The authors’ common-set analysis covers 121,680 valid trials across 12 models.

What did the authors report about failed attempts?

They report that 79,853 attempts in the common-set analysis failed executable checks. Among those failures, 67.24% had no final tool error despite a state-changing tool call. The authors also report incorrect field values, unintended extra effects and missing required effects; those categories overlap.

Do the reported scores predict performance in a company’s systems?

Not by themselves. The scores describe outcomes in the benchmark’s tested setup. The supplied material does not show that they predict performance across live systems, different policies or long-term use; organizations would need to test workflows relevant to their own operations.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role Of Anthropic’s Hardware Standard In AI Progress

Anthropic has launched a limited preview of its Model Hardware Standard, aiming to streamline AI integration with physical equipment in labs and factories.

5 Things To Know About Anthropic’s Claude Code Cloud Credits For Pro And Max

Anthropic is offering Pro and Max subscribers up to $250 in Claude Code cloud credits, but the tier amounts and terms remain unclear.

Ubisoft Begin Removing Ubisoft Connect From Steam With The New Ubisoft Connect Services

Ubisoft is gradually removing Ubisoft Connect from Steam as part of the rollout of the new Ubisoft Connect S service, causing concern among gamers and developers.

RavynOS: Pre-alpha Open-source OS Based On Darwin, FreeBSD, Apple Open-source

A new pre-alpha open-source operating system called RavynOS has emerged, built on Darwin, FreeBSD, and Apple open-source components, sparking increased interest.