DiagEvo Paper: Self-Evolution via Hierarchical Error Memory Beats R-Zero by 4.5 Points on Qwen3-8B Math
Summary
DiagEvo trains a diagnostician module to cluster solver failures into a hierarchical skill-node memory, then hands the causes to a challenger that generates targeted intermediate-difficulty questions with double-confidence filtering. On Qwen3-8B the framework hits 72.3% mean accuracy across five math benchmarks — 4.5 points above R-Zero — and averages 57.4% over nine benchmarks total, 1.1 points above DARC. The paper argues self-play stalls when the solver has no signal about which errors it keeps repeating.
Originally reported by huggingface.co
Read the original article →Original headline: DiagEvo Paper: Self-Evolution via Hierarchical Error Memory Beats R-Zero by 4.5 Points on Qwen3-8B Math