StepAudio 3 Realtime scores 98.9 on Artificial Analysis Full-Duplex Bench

Found first: a primary source the press has not covered yet.

StepFun published a technical report on September 12 for StepAudio 3 Realtime, a real-time conversational audio model that scores 98.9 overall on the Artificial Analysis Full-Duplex Benchmark. The paper lists roughly 90 authors, with Bin Lin among the leads, and was submitted to arXiv four days before the model's API release on September 15, 2026.

What the source says

The architecture rests on three stated components: Deep Perception, which captures acoustic cues to interpret user intent; Seamless Duplex, which models synchronized audio streams to handle pauses, backchannels, and interruptions; and Think-While-Speaking, which runs private chain-of-thought reasoning in parallel with spoken output rather than sequentially. The paper also reports 90.6 on MMSU and 56.0% macro task-success on the τ-Voice benchmark. StepAudio 3 Realtime is part of a broader StepAudio 3 family that became available through the StepFun API on September 15, 2026.

Why it matters

The Artificial Analysis Full-Duplex Benchmark is the main public index for real-time conversational voice AI, covering latency, turn-taking, and interruption handling. A 98.9 overall score on that index is a strong result and, at the time of submission, sits at the top of the published leaderboard. The Think-While-Speaking design is the core technical claim: parallel reasoning during speech would remove one of the main latency bottlenecks in voice agents, where current systems typically pause generation to compute before speaking. StepAudio 3 Realtime is in production as of September 15, available via the StepFun API.