Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Microsoft and Hugging Face Release ThinkingBox, a Benchmark That Judges Agents by Database State

Microsoft and Hugging Face have released ThinkingBox, a benchmark that grades AI agents on the database state and side effects they leave behind instead of the answers they produce. Across 507 business workflows run twenty times each, it finds that many agents finish cleanly while still writing the wrong records.

Published
微软与 Hugging Face 发布 ThinkingBox:用数据库状态而非回答检验 AI 智能体
Image source: huggingface.co

Microsoft and Hugging Face have released ThinkingBox, a new benchmark that grades AI agents on the records they leave behind rather than on the sentences they generate. The project is now available through Hugging Face, and the two organizations describe it as a way to test whether an agent can complete a task not once, but twenty times in a row.

ThinkingBox runs an agent against isolated MCP tool sessions and then inspects the terminal backend state and side effects it produces. It covers 507 stateful business workflows, each executed 20 independent times from an identical clean backend against a range of LLM models, so a single lucky pass cannot stand in for dependability.

The benchmark is illustrated with a support case. In a task adapted from the suite, an agent handling a delayed appliance delivery makes nine well-formed tool calls and reads the refund policy correctly, then closes the ticket as resolved. The carrier exception is still open, so the required end state was on hold, and the customer never receives a real answer.

That mismatch is the point. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error, while executable checks found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%.

The project argues that a tool call is not an outcome. Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. As the team puts it, a trajectory is a claim, database state is the evidence, and repetition is the trust test.

To separate breadth from consistency, ThinkingBox reports three numbers: pass@1 for how a model usually does, pass@20 for whether it can ever solve a task, and observed 20/20 for tasks that passed all twenty recorded attempts. The last metric is used as a literal count, with no estimator or smoothing, which the authors say is the column that matters for work touching real records.

The implication is that most leaderboards publish single-attempt scores, yet a high score once does not guarantee the same behavior on repeat. For teams wiring agents into customer support, refunds, insurance, or banking workflows, the gap between what a model can do and what it does every time is the operational risk.

ThinkingBox can be run through OpenEnv, and the full methodology is laid out in the accompanying ThinkingBox paper. The blog is a joint effort by Microsoft and Hugging Face, with co-authors and reviewers from several universities. The open question it leaves teams with is how to weigh consistency, not just peak accuracy, when an agent is allowed to change live records.

Why it matters

By grading database state rather than the answer, ThinkingBox pushes agent evaluation toward dependability, a shift that matters for any team letting agents modify real records.

MicrosoftHugging FaceAgentBenchmark
Back to realtime news

Nearby Updates

All