Microsoft benchmark grades agent runs by database state, not tool calls
ThinkingBox shifts agent evaluation to terminal backend state; the gap between one-off success and 20-for-20 consistency is the number to watch.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Microsoft and Hugging Face released ThinkingBox, a benchmark that grades agents on the terminal database state they leave behind, across 507 business workflows run 20 times each.
Across 121,680 valid trials on 12 models, 79,853 failed the executable checks; 67.24% of those had clean tool calls and no reported errors — the failures were wrong field values (77.61%) or unintended side effects (43.30%).
The consistency gap is larger: Kimi-K3 solves 93.89% of tasks at least once but only 13.41% on all 20 attempts; Claude Opus 5 solves 79.09% at least once and 47.53% every time.
The results are Microsoft and Hugging Face's own first-party measurements; the benchmark and dataset are open-sourced and reproducible via OpenEnv.