RealSWE: 88% of Real Coding Requests Are Absent From SWE-Bench Format, Cutting Agent Scores 6.4pp
Summary
Quantifies the distribution mismatch between SWE-bench design (structured, formal, information-rich) and real user requests (short, casual, minimal context): realistic inputs cut agent resolution rates 6.4pp on average and can flip model rankings, meaning current benchmark scores systematically overstate real-world coding-agent performance.
Originally reported by paper
Read the original article →Original headline: RealSWE: 88% of Real Coding Requests Are Absent From SWE-Bench Format, Cutting Agent Scores 6.4pp