Realtime AI News
Microsoft and Hugging Face Release ThinkingBox, a Benchmark That Judges Agents by Database State
Microsoft and Hugging Face have released ThinkingBox, a benchmark that grades AI agents on the database state and side effects they leave behind instead of the answers they produce. Across 507 business workflows run twenty times each, it finds that many agents finish cleanly while still writing the wrong records.

Microsoft and Hugging Face have released ThinkingBox, a new benchmark that grades AI agents on the records they leave behind rather than on the sentences they generate. The project is now available through Hugging Face, and the two organizations describe it as a way to test whether an agent can complete a task not once, but twenty times in a row.
ThinkingBox runs an agent against isolated MCP tool sessions and then inspects the terminal backend state and side effects it produces. It covers 507 stateful business workflows, each executed 20 independent times from an identical clean backend against a range of LLM models, so a single lucky pass cannot stand in for dependability.
The benchmark is illustrated with a support case. In a task adapted from the suite, an agent handling a delayed appliance delivery makes nine well-formed tool calls and reads the refund policy correctly, then closes the ticket as resolved. The carrier exception is still open, so the required end state was on hold, and the customer never receives a real answer.
That mismatch is the point. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error, while executable checks found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%.
The project argues that a tool call is not an outcome. Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. As the team puts it, a trajectory is a claim, database state is the evidence, and repetition is the trust test.
To separate breadth from consistency, ThinkingBox reports three numbers: pass@1 for how a model usually does, pass@20 for whether it can ever solve a task, and observed 20/20 for tasks that passed all twenty recorded attempts. The last metric is used as a literal count, with no estimator or smoothing, which the authors say is the column that matters for work touching real records.
The implication is that most leaderboards publish single-attempt scores, yet a high score once does not guarantee the same behavior on repeat. For teams wiring agents into customer support, refunds, insurance, or banking workflows, the gap between what a model can do and what it does every time is the operational risk.
ThinkingBox can be run through OpenEnv, and the full methodology is laid out in the accompanying ThinkingBox paper. The blog is a joint effort by Microsoft and Hugging Face, with co-authors and reviewers from several universities. The open question it leaves teams with is how to weigh consistency, not just peak accuracy, when an agent is allowed to change live records.
Why it matters
By grading database state rather than the answer, ThinkingBox pushes agent evaluation toward dependability, a shift that matters for any team letting agents modify real records.
Nearby Updates
All10/04, 05:36
Amazon Tied to $8B Nvidia AI Chip Sale-Leaseback Deal
Amazon is linked to a roughly $8 billion sale-leaseback deal involving Nvidia AI chips, in which it would sell the chips and lease them back for continued use. The move highlights how cloud providers are turning pricey AI compute into financeable assets as the GPU arms race strains balance sheets and capital budgets.
10/04, 05:00
NASA Uses AI to Read the Moon Before Humans Return to Its Surface
NASA is using artificial intelligence to study the Moon ahead of the return of human crews to its surface, according to a report from ColombiaOne. The story describes AI helping to interpret lunar data before future missions.
10/04, 02:50
GPT-6 Astra deciphers a 217-year-old secret letter from Napoleon to Marmont
GPT-6 Astra has helped reveal a 217-year-old secret letter written by Napoleon to Marmont, according to a report by Pasquale Pillitteri. The case is presented as a demonstration of how a next-generation model can be applied to historical document analysis.
10/04, 02:50
Google to block Gemini Flash for free users and pull Gemini Pro from AI Plus on October 9
Google plans to cut off Gemini Flash for free users and remove Gemini Pro from its AI Plus tier starting October 9, according to a report by Pasquale Pillitteri. The change redraws the boundary between free and paid model access and pushes stronger capabilities further up the subscription ladder.