ThinkingBox: when an AI agent says “done”, check the database
Microsoft and Hugging Face’s ThinkingBox tests whether agents leave a business workflow in the right state. It is a useful lesson for anyone building agents that can change real records.
By George the bot
Edited and approved by Faysal Aziz
Published
Updated

The agent sounded certain. It looked up an order, checked a late delivery, read the policy, opened a support ticket and told the customer the query was resolved. Yet the parcel was still stuck with the courier. The ticket should have remained on hold; the agent marked it solved.
That example, from Microsoft and Hugging Face’s ThinkingBox announcement, captures a familiar problem: a convincing final message is not evidence that the underlying job was completed correctly. If an agent can update customer records, make bookings or process a request, the state it leaves behind matters more than its closing sentence.
What ThinkingBox measures
ThinkingBox is a sandbox for stateful business workflows. ThinkingBox-Bench is its public benchmark of 507 synthetic tasks across retail, auto insurance, travel, neobanking and consulting. Each task gives an agent a user goal, a starting backend state, relevant tools and policies. After the agent acts, executable checks examine the final records and side effects.
That is different from merely checking whether a tool call was well-formed. An agent might call the correct ticketing tool but set the wrong status, modify an extra record or omit a required action. ThinkingBox can reject that outcome even if the agent’s prose sounds excellent. Most tasks are graded on state alone; a smaller set also needs a response check for things a database cannot capture.
The authors run each task 20 times from fresh, isolated starting states. That distinction is important. One successful attempt shows that an agent can sometimes do the job; repeated success is a better clue to whether the workflow is dependable.
The numbers need careful reading
In the published results, even the leading model’s overall single-attempt success estimate was 67.16%. The authors also report how many tasks passed on all 20 observed attempts. For that model, it was 241 of 507 tasks, or 47.53%. These are benchmark results, not a prediction for your own agent or business process.
ThinkingBox also reports that, in a common-set analysis, many failed attempts ended cleanly, made a state-changing tool call and had no final tool error. The problem was not always an obvious crash. Often, the system reached a plausible ending with the wrong state. The benchmark’s failure categories are observable diagnostics, not proof of a single underlying cause.
There is a naming trap here. In the post, “pass@20” means a task succeeded at least once in 20 tries. That measures breadth, not reliability. The stricter measure is the share of tasks that passed all 20 observed runs. If you are evaluating an agent that changes records, look at the latter alongside the ordinary per-attempt success rate.
Why this matters outside the benchmark
Imagine an appointment agent. Its final reply says, “Your booking is confirmed.” Did it create the booking? Was it for the right person and time? Did it cancel a different appointment by accident? A transcript check alone may miss those questions.
ThinkingBox’s useful idea is to define success in terms of the final, observable business state. That can help with support tickets, refunds, travel changes and internal back-office work. It does not mean every company should install this benchmark: its OpenEnv adapter currently needs externally managed proxy and tool services, a search service, and model endpoints. The released adapter is designed for evaluation, not a ready-made business agent.
Nor does a 20-out-of-20 result guarantee production reliability. The tasks are synthetic; your policies, tools, customers and failure modes may differ. Treat the benchmark as a way to ask better questions, not as a certificate of safety.
What you could learn and try
You can borrow the testing principle without reproducing the full harness. Pick a low-stakes workflow in a staging environment and define four things before testing the agent:
- Starting state: the exact records and permissions the agent should see.
- Required state: what must be true when the job is finished.
- Forbidden side effects: records or fields that must remain untouched.
- Repeated trials: several fresh runs of the same case, including awkward variants and tool failures.
For a support ticket, an assertion could require status “on hold” until a carrier exception clears. Another could check that exactly one ticket was created for the order. Record the agent’s final message too, but do not let “I’ve done it” stand in for those checks.
This is a useful skill set beyond ThinkingBox: state-based assertions, isolated test fixtures, repeatability, and distinguishing a failed tool from an incorrect business decision. When an agent fails, the trace helps diagnose what happened; the final state tells you whether the job was actually done.
The takeaway
An agent is not reliable because it speaks confidently, calls tools or succeeds once. The real test is whether it leaves the right records behind, avoids unwanted changes and can do that again. ThinkingBox makes that gap visible. Your own workflow tests should too.