← October 9, 2026 briefing · added in the 12:00 KST update
Microsoft researchers release 'ThinkingBox', a benchmark testing whether agents succeed at the same task 20 times in a row (community post)
What happened
A paper author identifying as a Microsoft employee introduced the agent benchmark 'ThinkingBox-Bench' on Reddit's r/MachineLearning. The benchmark consists of 507 policy-based business workflows across five domains: retail, travel and hospitality, auto insurance, internal IT at a neobank, and consulting IT/HR. It grades not on whether the agent says it finished the task, but on whether the database is actually in the correct state after the task ends. It also compares whether an agent succeeds once versus succeeding all 20 times across 20 repeated runs. According to the author, the paper, code, and dataset are public and also available on Hugging Face OpenEnv, but the original evaluation trajectories have not been released.
Why it matters
When companies use agents for real work, consistency across repeated runs and whether the actual system state changed correctly matter more than a single success rate. As a public evaluation tool that measures this gap, it can serve as a reference when assessing agent adoption. Specific result figures should be checked in the original paper.
Confidence medium · official source pending
Sources
More about Microsoft
-
Microsoft releases a version of its coding model MAI-Code-1.1-Flash optimized to run locally on PCs, timed with the Surface Laptop Ultra launch (report)
According to AI Times, on October 7 (local time) Microsoft released a version of its coding model 'MAI-Code-1.1-Flash' optimized to run...
-
Microsoft prices RTX Spark-based Surface Laptop Ultra from $2,599 and opens $5,999 Surface RTX Spark Dev Box preorders (report)
Specs and pricing are now public for the Surface Laptop Ultra, which uses NVIDIA's Arm-based RTX Spark chip. It starts at $2,599 with an...
-
Microsoft and NVIDIA unveil RTX Spark-based Surface devices; Windows Copilot gets local file access and OS-wide actions
Microsoft held a Windows and Surface event in San Francisco. On stage, NVIDIA CEO Jensen Huang and Microsoft CEO Satya Nadella said the...
-
GitHub makes Copilot local sandboxing generally available, adds local model discovery in the CLI and a dedicated leaked-secret detection model
GitHub announced three updates in its changelog. - Local sandboxing generally available (GA): local sandboxing for GitHub Copilot is now...
-
Microsoft Research releases Agent Lightning v1.0, a lightweight framework for reinforcement learning on existing agents
Microsoft Research has released Agent Lightning v1.0, an agent reinforcement learning framework of about 3,500 lines of code. It aims to...