Scale AI releases SWE-Bench Pro V2 with locked evaluations
Scale Labs released a 642-task version of its software-engineering benchmark with network-locked agent runs and pristine-image re-grading intended to make public results harder to game.
Scale AI released SWE-Bench Pro V2, cutting the public coding-agent benchmark from 731 tasks to 642 validated tasks across 11 repositories and updating its leaderboard. Scale Labs says it removed 89 tasks it considered invalid and introduced controls intended to make evaluation results more reliable and harder to game.
The main procedural change is a locked evaluation protocol. During the agent phase, general network access is disabled, only the model endpoint is allowed, and web-fetch tools are turned off. When an agent finishes, the system captures its code diff and replays the patch on a pristine image. That keeps the verifier out of the sandbox the agent modified. Scale says evaluators should report the re-graded score while publishing both the in-place and re-graded results.
Scale also rebuilt task environments from sanitized repository bundles intended to exclude fixing commits, stray references, stashes and hooks. Its repository includes probes that check whether a sandbox can reach the network or expose Git history. The company says it rewrote 529 problem statements, revised 214 test patches and 38 reference patches, and fixed dependencies in 211 images.
The dataset release makes V2 the default configuration while preserving the original 731-task version as the v1 configuration and v1.0 tag. The release also includes a 51-task HARD subset drawn from tasks failed by at least two of five model families under the locked protocol. Three tasks later judged ambiguous were excluded.
Scale’s published release-gate record says reference patches resolved all 642 tasks and empty patches resolved none. Those are company-run checks, not an independent reproduction. Scale’s broader claim that the new controls make public results more trustworthy has not been independently validated.
The refresh follows an independent September 8 preprint that identified reward-hacking opportunities and task-quality problems in the earlier benchmark, including misleading descriptions and improperly scoped tests. The research supports the case for revising the benchmark but does not validate V2’s controls. Scale also acknowledges that a locked runtime cannot remove information a model may already have encountered during training, including public repositories, fixing commits or the benchmark itself.
More news

MIT and Sakana AI propose SIFT to prioritize coding-agent benchmark runs

Together AI finds GLM-5.3 Flash 17x cheaper on DeepSWE

Meta launches Muse Code coding agent and upgraded Muse Spark 1.2 model

xAI launches Grok 4.5, undercutting Anthropic on coding-agent price
Dmytro Spodarets·Jul 9, 2026