QuoteBench: Coding Agent Success Rates Fall 55–73pp When Output Passes Through a Serializer
Summary
Current coding-agent benchmarks reporting matched execution scores cannot distinguish model failures from interface-layer failures—meaning practitioners are likely systematically overestimating agent capability whenever a serializer, wrapper, or shell reparser sits between the model and the terminal.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: QuoteBench: Coding Agent Success Rates Fall 55–73pp When Output Passes Through a Serializer