Categories:
Strategy
ai-agents benchmark evaluation microsoft buying-guide

Your Agent Demo Worked Once. Ask the Vendor for a 20/20 Score Before Signing.

Feature image for Your Agent Demo Worked Once. Ask the Vendor for a 20/20 Score Before Signing.

A customer’s $745 appliance had been sitting in a courier exception for fifteen days. They wrote in to get it sorted.

The AI support agent that picked up the ticket did almost everything right. Nine tool calls, each one sensible. It pulled up the order, documented the case, read the refund policy correctly. Then it closed the ticket as “resolved” when the workflow required the status “on hold.” The customer never got an actual answer. The agent reported success the entire time.

That ticket opens ThinkingBox, a benchmark Microsoft built with Hugging Face that grades AI agents on the records they leave behind instead of the sentences they produce. If you are evaluating agent vendors this year, the numbers in it should change how you run your next procurement call.

The words check out. The database does not.

Most agent benchmarks stop at the conversation. Did the model call the right tool? Did the final answer sound correct? ThinkingBox goes one step further and inspects the backend state after the agent finishes: which fields changed, what side effects fired, whether the required end state was actually reached.

That distinction sounds academic until you see the failure data. The team ran 121,680 valid trials across 12 models. About 80,000 of those attempts failed the executable checks. Here is the part that should worry anyone running agents in production: 67% of those failures ended cleanly. No errors. A state-changing tool fired. The agent confidently reported it had resolved the thing.

And the wrongness was specific. Of those clean-looking failures, 77.6% wrote wrong field values, 43.3% triggered unintended side effects, and 25.4% skipped required effects entirely. An agent can read the refund policy perfectly, quote it back to you, and still stamp the wrong status on the ticket. That is not a hypothetical. It is the modal failure.

Why one successful run means almost nothing

The second design choice in ThinkingBox is repetition. Every task runs 20 times from an identical clean backend, and the headline metric is “observed 20/20”: tasks that passed all twenty attempts.

This is the difference between “it can do this” and “you can depend on it.” A model with a 90% per-attempt success rate, which would sit near the top of most leaderboards, fails to string together twenty clean runs about 88% of the time. Run the arithmetic yourself: 0.9 to the 20th power is roughly 0.12. Demo-day success is a coin flip dressed up as a capability claim. For workflows that touch refunds, claims, or account changes, you need the boring version of the number.

The benchmark covers 507 stateful workflows across retail, insurance, travel, neobanking, and consulting. It runs through OpenEnv, so you can execute it yourself instead of trusting a vendor’s slide deck. The paper is on arXiv (2608.19741), and the code is open.

How to pressure-test an agent vendor

Take these four questions into your next procurement conversation:

  1. Ask for consistency scores, not pass@1. “What fraction of tasks pass twenty runs out of twenty?” If the vendor quotes a single-attempt number and cannot produce a repeated-run score, they have not measured reliability. They have measured luck on a good day.
  2. Ask how success is verified. If the answer involves reading the agent’s final message, that is the exact failure mode ThinkingBox just quantified. You want executable checks against the end state of the actual system of record.
  3. Run a pilot on your own workflows, scored on state. Give the agent ten of your real tickets, orders, or claims. Check the database afterward, not the chat transcript. The appliance ticket from the opening of this post would have passed a transcript review with flying colors.
  4. Watch for silent side effects. Two fifths of the clean failures in the study triggered changes nobody asked for. Your test plan needs to diff the whole record, not just the fields the agent was supposed to touch.

If you are building the agent yourself

The same rules apply internally. A passing demo in a sprint review tells you the pipeline works, not that the agent is dependable. Agent workflows that touch production data deserve the same treatment as database migrations: repeated trials, state diffs after every run, and a hard gate before anything ships. Treat a workflow that passes nineteen times out of twenty as a failed build, because at volume that is exactly what it is.

What to do now

Read the Hugging Face post first. The ticket example alone is worth five minutes of your time, and it is the fastest way to get your team aligned on why transcript reviews miss real failures. Then pull the benchmark through OpenEnv and point it at a workflow you actually run. If your current agent scores 20/20 on it, you have something you can lean on. If it does not, you just found out in a sandbox instead of in your ticket queue, with a real customer’s $745 order on the line.

Related Articles