Nine Frontier LLMs Average 26.7% Pass Rate on Auto-Built Real-Deployment Agent Benchmarks

Found first: a primary source the press has not covered yet.

Researchers at ByteDance and the University of Illinois at Chicago built an automated system that converts real-world agent deployment traces into targeted behavior benchmarks, then ran nine frontier LLMs against them and found a mean pass rate of 26.7%. The paper, TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces, was submitted to arXiv on September 27, 2026.

What the source says

TraceDance drew from 252,557 deployment sessions to produce 107 benchmarks containing 4,125 instances. The pipeline has two components: Anchor-and-Confirm, which combines programmable retrieval with candidate-level verification, and an Anchor Synthesis Loop, which generates and refines behavior specifications for custom requirements. It fulfilled 95.3% of build-target requests. Human annotators confirmed the requested behavior was present in 84% of sampled instances, and automated grading agreement matched inter-annotator reliability levels. The nine frontier LLMs were evaluated at recorded decision points using behavior-specific rubrics, without requiring reference answers or environment simulation.

Why it matters

Standard agent benchmarks measure whether a task is completed. TraceDance measures whether the agent behaved correctly during execution, a distinction the authors make explicit: an agent can finish a task while exhibiting undesirable behavior along the way. The 26.7% mean pass rate across nine frontier models on benchmarks derived from real deployments puts a number on that compliance gap. Because the benchmarks are generated from actual deployment traces rather than hand-authored scenarios, they capture failure modes that fixed benchmark suites cannot anticipate. For teams shipping agents into production, that means the gap TraceDance measures is the one that shows up in real users' sessions.