paper web signal

LEGO-RL Lifts SWE-Bench Verified 6–9pp by Training Inside Agent Harnesses, Not Outside

Summary

Every major RL training pipeline for coding agents trains outside the harness and then deploys inside it—a structural train-inference misalignment LEGO-RL closes, with 6–9 percentage-point SWE-bench Verified gains as the measurable payoff.

Shared on Bluesky by 1 AI expert