UpShaqo
Intelligence desk
Agent Trust & Security Source-backed analysis

A New Benchmark Grades Agents on What They Actually Changed

Microsoft and Hugging Face built a benchmark that ignores what an AI agent says and checks only what it left behind in the database — and the gap between the two is larger than most teams assume.

UpShaqo Editorial IntelligenceOctober 4, 20266 min read
Intelligence standard

Independent UpShaqo analysis built from fresh, attributed sources. We explain the impact instead of repeating the announcement.

Read for leverage: focus on the workflow change, the customer problem, and the next action—not only the product announcement.

A customer's $745 kitchen appliance sat stuck in a courier "exception" at a Nashville distribution center, fifteen days late. An AI support agent went to work: it pulled the order, checked tracking, reviewed the customer's profile, read the refund policy twice, confirmed no ticket existed, opened one, and correctly determined her account segment didn't qualify for late-delivery compensation. Nine clean tool calls. Then it closed the ticket as resolved and told her, "Since your query is resolved, is there anything I may assist you with?"

The carrier exception was still open. The required end state was on hold, not solved. And the customer never got a real answer to what she actually asked. Every tool call looked correct. The database disagreed.

That gap — between an agent sounding finished and an agent actually finishing — is the subject of ThinkingBox, a new benchmark built by Microsoft and released through Hugging Face. It doesn't grade transcripts. It grades the records left behind.

What ThinkingBox Actually Checks

Most agent evaluations score whether a tool call was valid or whether the final response sounds right. ThinkingBox runs agents through 507 stateful business workflows — retail, auto insurance, travel, neobank, consulting — inside isolated MCP tool sessions, then inspects the terminal backend state after the agent stops talking. Did the ticket land in the right status? Did the refund post the correct amount? Did the agent create a side effect nobody asked for?

The scale of the mismatch is the headline finding. In a common-set test across 12 models and 121,680 valid trials, 79,853 attempts failed the executable state checks — and of those failures, 67.24% had terminated cleanly, called a state-changing tool, and reported no error. The agent believed it was done. The ledger said otherwise, with wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%.

A transcript is a claim. State is the evidence.

One Right Answer Isn't a Right System

The benchmark's second contribution is almost more consequential than the first: it runs every task 20 independent times from an identical clean state, and reports three separate numbers — pass@1 (how it usually does), pass@20 (can it ever do this at all), and observed 20/20 (can it always do this).

Those numbers tell very different stories about the same model. Kimi-K3 solves 93.89% of the benchmark at least once — the broadest coverage of any model tested — but passes only 13.41% of tasks on all 20 attempts. Claude Opus 5 solves fewer tasks at least once (79.09%) but completes 47.53% of the benchmark every single time.

Upgrades don't automatically buy dependability, either. Claude Opus 5.5 beats Claude Opus 5 on pass@1 (67.16% versus 66.50%) and solves more tasks at least once — but passes exactly the same 241 tasks on all 20 attempts. Half a point of headline accuracy bought zero additional consistency.

Analysis: For any workflow that writes to a real database — a refund, a policy change, a booking cancellation — pass@1 and pass@20 are the wrong columns to optimize. A model that can technically solve a task once in twenty tries isn't a feature; a model that solves it correctly every time it's asked is the actual product requirement.

The Price of Being Right Every Time

ThinkingBox also prices the gap. Researchers computed cost per successful task attempt (what it costs to get one right answer) and cost per dependable task (what it costs for a model to pass all 20 attempts), using list token rates.

The two rankings diverge sharply. GPT-5.6 Sol has the lowest cost per single success at $0.127, but costs $9.76 per dependable task — more expensive to run reliably than GPT-5.4's $6.80. Claude Opus 5 and Claude Opus 5.5 pass the identical 241 tasks consistently, but Opus 5 costs $13.30 per dependable task against Opus 5.5's $7.80 — meaning the newer model fully dominates the older one on cost without touching its reliability ceiling.

The cheapest way to get a right answer, in other words, is not the cheapest way to get a dependable one. Those are two different purchasing decisions, and most procurement conversations only ask about the first.

Where Agents Actually Break

The benchmark's failure analysis is the most actionable part for anyone building, not just evaluating, agents. Across the ablation study, roughly four in five failures trace to tool usage problems — 79.9% — rather than reasoning errors. Wrong state updates account for 10.3%, incomplete resolutions 7.0%, and failure to act at all just 2.9%.

The practical pattern: agents usually get far enough into a workflow, then fail to recover from a tool error, a failed precondition, or an empty lookup. That's a retry-and-error-handling problem sitting in front of a model-capability problem. Domain difficulty compounds it — retail workflows average 59.52% pass@1 across tested models, while auto insurance averages just 33.83%.

A Scenario: Staffing a Refund Desk

Consider an operations leader at a mid-size retailer deciding which model to put behind an agent that processes refund exceptions — the exact category of task in the Nashville example. Three candidates surface from the benchmark data:

  • Kimi-K3 solves the broadest range of edge cases and leads outright on retail tasks at 82.24% pass@1, useful if the team wants an agent that can attempt almost anything. But only 13.41% of tasks pass all 20 times, meaning roughly six in seven successful-looking runs carry some risk of silent recurrence on a repeat attempt.
  • Claude Opus 5.5 trails Kimi-K3 on raw retail coverage but ties for the highest observed 20/20 rate in the field at 47.53%, and does it at $7.80 per dependable task — cheaper than Claude Opus 5's identical reliability.
  • GPT-5.4 is the cheapest route to a dependable task at $6.80, though it clears the 20/20 bar on fewer total tasks (128) than either Claude model.

This is UpShaqo's inference, not a finding in the paper: a team automating irreversible actions — refunds, cancellations, account changes — should weight the 20/20 and cost-per-dependable-task columns over pass@1, even if that means passing on the model with the flashiest single-run score. A team using agents only for draft suggestions a human reviews can afford to optimize for breadth instead.

What to Check Before an Agent Touches a Ledger

The ThinkingBox authors offer a short, testable set of mitigations, explicitly noting they haven't yet measured the lift from any of them on this benchmark:

  • Treat the 20/20 rate as a design input, not a pass/fail verdict on a vendor.
  • Check terminal state before committing a change, not the model's summary of what it did.
  • Classify tool and system errors so retries target the ones that are actually recoverable.
  • Narrow the tool surface to only what the workflow needs.
  • Require human approval on changes that are expensive to reverse.

The benchmark itself is now available through Hugging Face via the OpenEnv interface, with a pinned data release and bundle hash so results are verifiable rather than asserted. For any team already running agents against production data, the more immediate move may be smaller than adopting a new benchmark: pull one recent transcript that ended in a clean-sounding closure, and check what actually changed in the record behind it.

Sources

#AI agents#Microsoft#Hugging Face#enterprise AI#LLM benchmarking#agent reliability

Two doors. Pick one.

Hire the team.
Or become it.