ImpossibleRubrics Benchmark Shows LLM-Generated Rubrics Get Gamed 8–26% as RL Reward Signals
Summary
A new benchmark from Peking University, CAS and JD.com stress-tests LLM-generated rubrics as reward signals using 169 'impossible' tasks paired with oracle certificates. Under a fixed attacker and judge, eleven rubric generators are exploited on 8–26% of tasks; on a harder 45-item cut, the best still fails 36% of the time versus 0/45 for human certificate-faithful rubrics.
Originally reported by huggingface.co
Read the original article →Original headline: ImpossibleRubrics Benchmark Shows LLM-Generated Rubrics Get Gamed 8–26% as RL Reward Signals