arxiv.org web signal

AgSpec Paper Clocks 4.76x Decoding Speedup for Coding Agents

TL;DR

  • AgSpec reports up to 4.37x generation throughput at batch size 1 and 4.76x at batch size 16 versus autoregressive decoding.
  • The method retrieves draft tokens from session, workspace, and global corpora and indexes opened files in the agent's emission format.
  • It outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings on two repository-level multi-agent coding benchmarks.

A paper posted to arXiv on 1 October reports a retrieval-based speculative decoding system that raises generation throughput over autoregressive decoding up to 4.37x at batch size 1 and 4.76x at batch size 16 on two repository-level multi-agent coding benchmarks.

The system is called AgSpec. Authors Sumin Lee, Sukmin Cho, Suengjae Lim, and Youngjin Kwon frame the problem bluntly: "much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns."

Their fix has two parts. AgSpec "retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format." And it "bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback."

The paper reports that AgSpec "outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings" and "remains effective on benchmarks without a repository or a multi-agent pipeline." The abstract refers to the five retrieval drafters only as a count, and does not name the base models or specific benchmarks used.

It is one of several inference-side results moving through our coding tools coverage this week.