OFICIAL Hugging Face Blog

The Agent Said It Was Done. The Database Disagreed.

What happened
Based on Hugging Face Blog · Oct 03, 2026

Hugging Face introduces ThinkingBox, a benchmark grading AI agents by backend state changes rather than tool calls, revealing widespread consistency failures across 507 workflows and 12 models.

The Agent Said It Was Done. The Database Disagreed.
Hugging Face Blog — Hugging Face
Key points
·
Microsoft’s ThinkingBox grades AI agents by backend state changes, not tool calls or responses, across 507 workflows.
·
79,853 of 121,680 trials failed executable checks despite clean terminations and state-changing tool invocations.
·
Only three models retained over 70% of their initial pass@1 scores after 20 repeated trials in the benchmark.
Key numbers
·
The benchmark runs each agent 20 times across 507 stateful business workflows to test reliability, exposing gaps between claimed actions and actual database outcomes.
·
Across 121,680 trials involving 12 LLM models, 79,853 attempts failed executable checks despite clean terminations and state-changing tool invocations.
·
61% involved wrong field values, 43.

Microsoft’s ThinkingBox evaluates AI agents by examining the terminal backend state and side effects they leave behind, not by their generated responses or tool calls. The benchmark runs each agent 20 times across 507 stateful business workflows to test reliability, exposing gaps between claimed actions and actual database outcomes. The approach isolates agents in controlled MCP tool sessions, grading them on whether the backend reflects the required end state after each run.

A real-world example illustrates the problem: an AI agent correctly performed nine tool calls to address a delayed appliance delivery, but the database showed the ticket status as 'solved' while the required state remained 'on hold.' The agent’s final response claimed resolution, yet the underlying issue persisted. ThinkingBox flags such discrepancies by comparing the agent’s tool calls with the actual backend state, highlighting the difference between procedural correctness and outcome accuracy.

Across 121,680 trials involving 12 LLM models, 79,853 attempts failed executable checks despite clean terminations and state-changing tool invocations. Of these failures, 77.61% involved wrong field values, 43.30% had unintended side effects, and 25.36% missed required effects. The benchmark’s repetition requirement—20 identical runs per task—reveals that models performing well once often fail consistently, undermining trust in their reliability.

ThinkingBox is now available via Hugging Face, allowing users to replicate the benchmark using OpenEnv. The tool measures three key metrics: pass@1 (single-attempt success), observed 20/20 (tasks passing all 20 runs), and survival rate (how much of the pass@1 score holds after repetition). Results show wide variability, with only three models retaining over 70% of their initial pass@1 scores after 20 trials, while others drop to near 8%.

Original source → Deals on Clipraptor.com →