EvoRubric lets an LLM evolve its own grading rubrics in RL
TL;DR
- EvoRubric uses a shared policy as both Reasoner and Rubric Generator, with a frozen copy of the initial policy acting as a Meta-Verifier for its criteria.
- Across five Medical, Writing, and Science benchmarks it averages 56.28 at 8B and 61.13 at 14B, beating matched baselines by 3.09 and 2.06 points.
- It fuses criterion-validity feedback, Leave-One-Out peer consensus, and a persistent memory pool into rewards that jointly train both policy roles.
EvoRubric assigns an LLM two jobs at once in reinforcement learning: produce the answer, and produce the rubric that will grade it. A frozen copy of the initial policy sits alongside as a Meta-Verifier to check whether the criteria the model invents are actually valid.
In the authors' own words, the system "combines criterion-validity feedback, response discrimination, and peer agreement to learn rubrics for open-ended generation." A separate frozen Grader scores responses, and a Leave-One-Out peer-consensus signal plus a persistent memory pool are turned into rewards that train both policy roles at once.
The arxiv preprint reports five-benchmark averages of 56.28 at 8B and 61.13 at 14B across Medical, Writing, and Science tasks, which the authors say beats the strongest matched baselines by 3.09 and 2.06 points over three training seeds. The abstract gives no per-benchmark numbers and does not describe the compute cost of running Reasoner, Rubric Generator, Meta-Verifier, and Grader in the same loop.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation