Microsoft releases ThinkingBox agent benchmark through Hugging Face
ThinkingBox checks whether AI agents leave the required backend state and side effects, while repeated trials separate occasional success from reliable execution.
Microsoft has released its ThinkingBox agent-evaluation framework and a 507-task benchmark through Hugging Face, giving developers a way to test whether an agent completes a workflow instead of relying on its final response alone.
ThinkingBox runs agents in isolated, stateful environments and grades the terminal backend state and side effects. Microsoft says the design can catch errors an execution trace might miss: tool calls can appear valid even when required changes are missing, incorrect changes are made or extra changes remain. The framework can also use a simulated user for follow-up exchanges and records each execution trace.
ThinkingBox-Bench v1.0 contains 507 workflows across five fictional business domains: 98 retail and e-commerce tasks, 104 travel and hospitality tasks, 100 auto-insurance tasks, 104 neobank-support tasks, and 101 consulting IT and HR-support tasks. Each task defines an initial backend state, a user goal and simulated-user context, available tools, policy constraints, and executable checks. A run passes only if all required checks succeed.
To grade the final state, the scenario server starts with a fresh database, applies the task’s initial-state patch and replays the expected tool interactions to create a golden state. It then compares a stable hash of that database with the hash of the database left by the agent. Thirty tasks also include yes-or-no dialogue rubrics evaluated by an LLM judge. The release gives no partial credit when an applicable state or dialogue check fails.
In its published evaluation, Microsoft ran every task 20 times. The aggregation tooling reports pass@1, pass@20 and pass^20: single-run success, at least one success in 20 attempts, and success on all 20 attempts, respectively. Microsoft’s benchmark explanation uses the last two measures to distinguish finding a successful trajectory through retries from repeating that success reliably. The results and failure analysis have not been independently reproduced across the full benchmark.
For context, Microsoft has also added agent evaluation and optimization tools to Foundry. ThinkingBox’s published design centers on controlled workflows and expected terminal state.
Microsoft maintains the ThinkingBox framework and benchmark data in separate repositories. Hugging Face provides a browsable dataset view, while Microsoft directs users to the tagged GitHub release to run the executable benchmark. The v1.0 documentation says the benchmark is for evaluation only and should not be used for prompt optimization, fine-tuning, reinforcement learning, reward-model training or other model optimization.
More news

Microsoft adds voice agents and optimization tools to Foundry

Microsoft relaunches Copilot with Home, Code and Autopilot

NetApp puts approval gates around Workload Factory AI agents
