DarwinX Evolves Agent Harnesses to 93.0% on WebArena-Infinity, Model Weights Frozen

Found first: a primary source the press has not covered yet.

A paper posted to arxiv on July 31, 2026 describes DarwinX, a system that applies population-based natural selection to agent harnesses while leaving the underlying language model entirely frozen. On WebArena-Infinity, audit-clean pass@1 rises from 43.5% to 93.0%.

What the source says

DarwinX, from a team of twelve including Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, and Zeyuan Chen, evolves the full agent scaffold: prompts, tools, skills, and control flow. Fitness is scored by each benchmark's own verifier, requiring no gold-standard solutions. An archive of alternative lineages is maintained to allow recombination across generations. On Terminal-Bench 2.1, the approach reaches 83.2% (+7.7 points) with a matched base model and 84.7% with a stronger one; TerminalWorld's held-out split comes in at 68.3%. A harness evolved on Terminal-Bench 2.1 transfers to SWE-bench Verified without modification.

Why it matters

The WebArena-Infinity figure, if it holds under independent replication, is the most significant claim: more than doubling a hard open-ended web-agent benchmark score without touching model weights shifts the question of where capability gains come from. The SWE-bench transfer adds a separate data point, a harness evolved on one benchmark applying directly to another without retuning. The paper claims "what evolves is general agent competence, not benchmark-specific patches," though a 93.0% pass@1 result on a benchmark where the prior state of the art sat below 50% warrants scrutiny before that framing is accepted. The reported average per-loop gain across four benchmarks is approximately 17 percentage points.