paper web signal

TraceDance auto-builds behavior tests; frontier LLMs pass 26.7%

TL;DR

  • TraceDance mines 252,557 deployment sessions to auto-build 107 behavior-specific benchmarks containing 4,125 test instances.
  • Nine frontier LLMs achieve a mean pass rate of only 26.7% at the paper's recorded decision points.
  • The pipeline fulfills 95.3% of build-target requests, and human annotators confirm the requested behavior in 84% of sampled instances.

Nine frontier language models managed a mean pass rate of only 26.7% on the TraceDance benchmark suite, a set of behavior tests built automatically from 252,557 deployment sessions.

The paper's premise is that an agent can finish a task while doing something you did not want along the way, and that fixed suites miss those cases. "An agent can complete a task while exhibiting undesirable behavior during execution," the authors write. "Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites."

TraceDance turns the trace pile into 107 behavior-specific benchmarks with 4,125 instances, "fulfilling 95.3% of build-target requests." The evaluation uses "decision-point continuation": the model has to produce the next turn at a recorded decision point, judged against a behavior-specific rubric, without a reference answer or environment replay. Human annotators confirm the requested behavior in 84% of sampled instances, and the paper reports that the automated grader's agreement with human pass/fail judgments is "comparable to that between the annotators."

The authors pitch the pipeline as feedstock for "the recursive self-improvement (RSI) loop," by "turning deployment problems into targeted benchmarks."

The abstract names neither the nine models tested nor per-model scores, and its experiments cover only coding and general tool use.