paper web signal

EvoDuet lifts Gemini-3.8-Flash 21 points via query co-evolution

TL;DR

  • EvoDuet is a bilevel loop that co-evolves candidate solutions with the search queries used to fetch supporting evidence, keeping model parameters frozen.
  • On Gemini-3.8-Flash, normalized discovery gain across 21 optimization tasks climbs from 61.3% to 82.3%; on GPT-5.6-Luna it moves 74.1% to 78.0%.
  • The method beats previously reported best scores on eight tasks including Swap Reduction on Q20 and Rosetta, but Qwen3.5-9B shows no benefit.

On Gemini-3.8-Flash, co-evolving the search query alongside the candidate solution lifts OpenEvolve's normalized discovery gain from 61.3% to 82.3%. On GPT-5.6-Luna the same loop moves the number from 74.1% to 78.0%. On Qwen3.5-9B it does nothing.

That spread is the headline finding of EvoDuet, an arXiv paper from eight authors led by Young-Jun Lee. The paper frames the problem flatly: "Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks." Bolting a web-search tool onto the loop is not enough, the authors argue, because the queries keep returning the same pages as the solutions change.

EvoDuet's response is a bilevel loop. Solutions and queries co-evolve while the model parameters stay frozen. At each iteration, "a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them." The queries fetch evidence from arXiv, GitHub code, and web docs, and the selected material feeds the next candidate.

Across 21 optimization tasks with one candidate per iteration, the best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta. The abstract does not publish per-task numbers, retrieval or API cost figures, or an explanation for why Qwen3.5-9B sits out the gains the two larger models enjoy.