LEGO-RL Lifts SWE-Bench Verified 6–9pp by Training Inside Agent Harnesses, Not Outside
Summary
Every major RL training pipeline for coding agents trains outside the harness and then deploys inside it—a structural train-inference misalignment LEGO-RL closes, with 6–9 percentage-point SWE-bench Verified gains as the measurable payoff.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: LEGO-RL Lifts SWE-Bench Verified 6–9pp by Training Inside Agent Harnesses, Not Outside