The agent said it was done. The database disagreed.

A customer writes in about a delayed kitchen appliance worth $745, stuck in a courier exception at a distribution center for fifteen days past its estimated delivery date. An AI support agent swings into action. It pulls the order, checks tracking, looks up the customer profile, reads the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and closes the ticket as resolved. Nine well-formed tool calls. A tidy reply: "Since your query is resolved, is there anything I may assist you with?"

Two things are wrong. The carrier exception is still open, so the required end state was a ticket on hold, pending resolution. And the customer never actually got an answer to her question. Under the way most AI evaluations are scored today, that run would be recorded as a success.

That gap is what a new benchmark called ThinkingBox sets out to measure. Published by Microsoft with Hugging Face on October 3, 2026, ThinkingBox grades AI agents not on the sentences they generate or the tool calls they make, but on the terminal database state and side effects they leave behind, then asks whether they can do it twenty times in a row. Its thesis, stated plainly in the paper's title: "One Success Isn't Reliability."

This article explains what the benchmark found, why one success is not reliability, and why the cheapest model per success is not the cheapest per dependable task. The caveats come first, because they matter: ThinkingBox is a vendor-run benchmark, built by Microsoft on its own grading framework, and every number below is vendor-reported rather than independently verified. A benchmark from a company that also sells AI models should be read as a well-documented argument, not settled science.

Grading by what was written, not what was said

Most agent evaluations, including leaderboard scores businesses use when choosing models, grade what an agent says or the calls it makes. An evaluator checks whether the final response sounds correct, or whether the sequence of tool calls looks valid. Both are proxies. Neither looks at what actually changed in the systems the agent was supposed to operate.

ThinkingBox flips that. It runs an agent against isolated tool sessions backed by a simulated business backend, then applies task-specific executable checks to the final state of the database. A trajectory is a claim. Database state is the evidence.

The check is concrete, not fuzzy. In the delayed-appliance example above, the executable check that fails is a single field: the ticket's status is "solved" where the required end state is "hold". That example is adapted from a benchmark task the authors identify as test_case_ST003_006, with the full trace in the paper's appendix.

The scope is large. ThinkingBox-bench contains 507 policy-conditioned workflows spanning retail and ecommerce, hospitality, auto insurance, neobank internal IT, and consulting IT and HR support. Every task is run twenty independent times, each from an identical clean backend, yielding more than ten thousand graded attempts per model.

The numbers: failures that look like wins

The headline finding comes from an ablation study across twelve models: 121,680 valid trials produced 79,853 attempts that failed the executable checks. Of those failures, 67.24 percent still terminated cleanly, invoked a state-changing tool, and reported no error. In other words, two-thirds of failures look like successes to any grader that watches the conversation rather than the database.

What did the executable checks find inside those clean-looking failures? Wrong field values in 77.61 percent of them. Unintended extra effects, records changed that should not have been, in 43.30 percent. Missing required effects, work that needed doing and was never done, in 25.36 percent. The categories overlap; a single failed run can be wrong in more than one way.

The failure-signature analysis is more actionable still: roughly four in five failures, 79.9 percent, are tool usage problems rather than reasoning problems. Wrong state updates account for 10.3 percent, incomplete user resolutions 7.0 percent, and no state-changing action at all 2.9 percent. If those shares hold beyond this benchmark, the practical implication is that most agent reliability work is plumbing, not philosophy: the models can usually think their way through the problem, but they fumble the mechanics of changing records correctly.

One success is not reliability

ThinkingBox reports three different numbers, and the difference between them is the whole story. Pass@1 is the share of all attempts that succeeded, the number most leaderboards publish. Pass@20 is the share of tasks solved at least once in twenty tries, a measure of breadth. Observed 20/20 is the share of tasks that actually passed all twenty recorded attempts, with no estimator and no smoothing, a measure of dependability.

The spread between the columns is dramatic. Kimi-K3, an open-weights model, has the broadest coverage of any model tested: it solves 93.89 percent of the benchmark at least once, 476 of 507 tasks, and leads retail outright at 82.24 percent pass@1. But only 13.41 percent of tasks, 68 of 507, succeed in all twenty attempts. A model that can do nearly everything once in five tries is not a model you can depend on.

Claude Opus 5 inverts the pattern. It solves fewer tasks at least once, 79.09 percent, but completes 47.53 percent of the benchmark, 241 tasks, on every single attempt. Its successor does not fix this. Claude Opus 5.5 scores higher on every-attempt average, 67.16 percent against 66.50 percent, and solves more tasks at least once, yet passes exactly the same 241 tasks on all twenty attempts. Half a point of headline accuracy bought no additional dependability at all.

The arithmetic makes the tradeoff stark: Kimi-K3 solves 75 more tasks at least once than Opus 5, while Opus 5 solves 173 more tasks consistently than Kimi-K3. For a business choosing a model for work that touches real records, the benchmark argues that pass@20 is the wrong column to look at. It is worth noting that a 20/20 pass on a twenty-trial sample is an observed count, not a proof of perfect reliability; a model could still fail on a twenty-first run. But the direction of the gap between pass@1 and 20/20 is hard to dismiss.

What consistency costs

The benchmark authors then ask a question most capability comparisons stop before: what does a dependable unit of work cost? They compute cost per successful task attempt by dividing a model's estimated cost for 507 attempts, priced at undiscounted list rates from a single OpenRouter endpoint per model, by the number of attempts that succeeded. They are careful to label this a comparative efficiency index, not an invoice and not the price of serving one production request.

On that measure, three models sit on the Pareto cost frontier, meaning no cheaper model matches their accuracy. GPT-5.6 Sol is cheapest per success at $0.127; GPT-5.4 raises accuracy by 3.45 percentage points for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success. Claude Opus 5 is dominated outright: at $0.475 per successful attempt and 66.50 percent pass@1, it is both more expensive and less accurate than Claude Opus 5.5.

Then they price consistency differently. Cost per dependable task divides the cost of the full twenty-run campaign by the number of tasks the model passed on all twenty attempts. GPT-5.4 is cheapest at $6.80 per dependable task, though only 128 tasks meet the bar. GPT-6 Astra reaches 231 dependable tasks at $7.45. Claude Opus 5.5 reaches 241 at $7.80. The ordering changes from the pass@1 ranking, and the cheapest model per single success, GPT-5.6 Sol at $0.127, costs $9.76 per dependable task.

The lesson for buyers is uncomfortable. A leaderboard score optimizes for how often a model is right. A business deploying an agent against live records needs to know how often it is right every time, and the price of that guarantee can differ by an order of magnitude from the price of one good answer. A purchasing decision made on pass@1 alone may be buying something that works once in five, at a per-dependable-task cost far above what the headline comparison implied.

Open, but vendor-run: reading the caveats

ThinkingBox is open. The benchmark dataset is released on Hugging Face under the CDLA-Permissive-2.0 license, the sandbox code is on GitHub under the Microsoft organization, and the harness is available through OpenEnv. The dataset viewer shows real task definitions: initial state patches for simulated backends, expected tool interactions, and executable checks, with a v1.0 release tag.

Several caveats belong next to any use of these numbers. The benchmark is vendor-run, from a company with a direct stake in agent deployment, and co-authored with Hugging Face, which also has commercial interests in the agent ecosystem. The pass@1 rankings in Table 2 are vendor-reported and were not independently re-run for this article. Model names in the tables include several this newsroom cannot independently confirm as released products, so the reader should treat the model list as reported by the authors. Cost figures use undiscounted list prices and exclude quantized endpoints, so real-world costs will vary. And a benchmark of simulated business workflows, however carefully constructed, is still a simulation; production environments have messier edges.

None of that blunts the central, verifiable methodological point, which does not depend on any particular model's score: grading agents by the database state they leave behind catches failures that response-level and tool-call-level grading misses, and 67.24 percent of failed attempts in this study would have been scored as clean by such graders. Any business buying agents on leaderboard scores should ask which column the score came from, and how many times each task was actually run.