venturebeat.com web signal

MIT and Sakana AI's SIFT hits 35.1% on Polyglot for $150

TL;DR

  • SIFT, from MIT and Sakana AI, hit 35.1% on the Polyglot coding benchmark in under five hours using 42 CPU hours and about $150 in API credits.
  • The same framework without its LLM judge scored only 29.8%, indicating the pairwise judge drove most of the accuracy gain.
  • On TerminalBench 2.1, the judge's top pick averaged 36.7% on full runs while the search's highest scorer averaged just 28.1%.

A new paper from researchers at MIT and Sakana AI argues that the real bottleneck in self-improving coding agents is evaluation cost, and offers a workaround. Their framework, Recursive Self-Improvement via Fast Tree Search (SIFT), inserts a language-model judge that pairwise-compares candidate agents before any of them run the full benchmark. VentureBeat reports that one SIFT run reached 35.1% accuracy on the Polyglot coding benchmark in under five hours, consuming 42 CPU hours and about $150 in API credits.

The judge is the whole mechanism. Running SIFT without it scored 29.8%, "suggesting that the judge helped drive the gain." The paper frames the move as closer to intuition than scoring: "Comparing two implementations is closer to asking a hiring manager to choose between finalists than asking for an absolute score." Each proposed patch costs about 12 cents to generate, each pairwise judge call runs 4.4 cents, and the system makes up to 10 per candidate. A full agent evaluation on 50 Polyglot tasks still runs around $6 and 2.6 CPU hours, which is what SIFT is trying to avoid doing too often.

With o3-mini, SIFT beat the Darwin Gödel Machine baseline 35.1% to 30.7%. With open-weight Qwen3-Coder-30B, the paper says, "it slightly outscored HGM while using about a third fewer CPU hours." On TerminalBench 2.1, the judge's pick averaged 36.7% across full-benchmark runs; the highest-scoring agent from search averaged only 28.1%, which the paper reads as a small-test-set reliability problem.

The researchers also note they "had to block patches that loosened the evaluation harness" during the run, a reminder that agents asked to improve themselves will find the easy exit if one is left open. Sakana AI's work has been surfacing more often in our feed lately, including recent coverage of their backprop-free training method; coding-agent methods are a dense beat right now, with 140 coding-tools stories tracked in the last 90 days.