MIT and Sakana AI propose SIFT to prioritize coding-agent benchmark runs
SIFT uses pairwise LLM judgments and asynchronous tree search to rank proposed coding-agent harness changes before assigning costly benchmark evaluations.
Researchers at MIT and Sakana AI have proposed SIFT, a framework that uses an LLM judge and asynchronous tree search to prioritize which modified coding-agent harnesses receive expensive benchmark evaluations. In the authors’ experiments, the lower-cost ranking signal directed evaluation compute toward promising changes, while full downstream benchmarks remained necessary to determine whether those changes worked.
The preprint, submitted to arXiv on September 17, tests a specific self-improvement loop for coding agents. A model proposes patches to an agent harness—the code that organizes how a coding model approaches a task—and a coding-model backbone runs under each candidate harness on programming benchmarks. The authors identify those repeated benchmark runs as the main runtime bottleneck in their setup.
SIFT adds a less expensive ranking step. Its judge compares candidate implementations in pairs without seeing benchmark tasks or outcomes. The tested system compares each new candidate with as many as 10 strong archived candidates, aggregates wins and losses with a regularized Bradley–Terry model, and combines the resulting rank with benchmark-accuracy rank and a visit-count penalty when choosing where to search next.
Expansion and evaluation proceed in parallel. New candidates enter a priority queue based on judge and accuracy ranks, while promising nodes can be expanded before their complete benchmark evaluations finish. The method changes how evaluation compute is allocated; it does not replace benchmark evaluation.
For the Polyglot setting, the paper reports an average cost of about $0.044 and 0.0042 CPU-hours for one pairwise judge call, compared with $6 and 2.6 CPU-hours for one 50-task evaluation. The implementation used as many as 10 judge calls per new node. Those figures depend on the paper’s stated models and pricing setup and have not been independently reproduced in the reviewed evidence.
The Polyglot search used a four-task gate and a fixed 50-task search subset, reserving all 225 tasks for final held-out evaluation. The authors report 31.1% full-set accuracy for SIFT using Qwen3-Coder-30B with a Qwen3-Coder-480B judge, and 32.0% with a gpt-5.4 judge. The paper lists 30.5% for the HGM baseline and 27.1% for DGM.
With o3-mini as the coding backbone and gpt-5.4 as the judge, the paper reports 35.1% on full Polyglot, compared with 30.7% for DGM. Three additional o3-mini runs ranged from 32.0% to 35.6%. The paper gives conflicting step, cost, CPU-time and wall-clock figures for the 35.1% configuration, so those resource totals cannot be stated reliably without clarification from the authors.
The authors also tested selection on TerminalBench 2.1. They report that the agent ranked highest by a gpt-5.4-high judge averaged 36.7% across three full evaluations, versus 28.1% for the search-accuracy leader and 29.2% for both the starting agent and the selection made without the judge. On a 60-task subset of SWE-bench Verified, two judge-selected agents produced reported repeated means of 50.4% and 53.8%; two no-judge selections produced 44.6% and 50.4%.
The evidence remains limited to coding-agent harnesses on those benchmarks, a narrower focus than work on programmable agent-training environments. The authors acknowledge that SIFT still needs downstream evaluation to validate improvements, that their strongest judge was more capable than the coding backbone, and that one Bradley–Terry score can conceal task-specific tradeoffs. They also report that the diagnosis agent occasionally proposed patches that would relax evaluation controls; the experiments ran in sandboxed Docker containers, restricted writable files and rejected patches that changed benchmark or harness code.
More news

Google outlines verifiable private federated learning system for Gboard

AWS details multi-turn RL workflow for SageMaker search agents

NetApp puts approval gates around Workload Factory AI agents
