StepAudio 3 Realtime posts 98.9 on Full-Duplex voice bench
TL;DR
- StepFun's StepAudio 3 Realtime reports 98.9 overall on the Artificial Analysis Full-Duplex Bench, 90.6 on MMSU, and 56.0% macro on τ-Voice.
- The architecture centers on Think-While-Speaking, running private reasoning in parallel with spoken output to hit reasoning targets in real time.
- On Artificial Analysis's public leaderboard, StepAudio 3 Realtime leads Conversational Dynamics ahead of Qwen Audio 3.0 Realtime Plus (98.4%) and GPT-Live-1 (97.3%).
StepFun's technical report for StepAudio 3 Realtime, posted to arXiv, claims a 98.9 overall score on the Artificial Analysis Full-Duplex Bench, alongside 90.6 on MMSU and a 56.0% macro task-success rate on τ-Voice.
The paper describes the system as an "audio-language foundation model organized around a continuous listen-converse-think-act loop." Its central trick, per the abstract, is Think-While-Speaking, "executing private reasoning in parallel with spoken delivery." In reasoning mode the model reports 73.0 macro on StepAudioChat.
On the Artificial Analysis leaderboard, StepAudio 3 Realtime leads Conversational Dynamics across 33 models at 98.9%, with Qwen Audio 3.0 Realtime Plus at 98.4% and GPT-Live-1 (Sol, low) at 97.3%.
StepFun said on X that Realtime "ranks #1 on Artificial Analysis for both Conversational Dynamics (98.9%) and Speech Reasoning (99.7%)," and that its ASR reaches a 1.7% word error rate. The company frames the launch as a five-model audio family spanning real-time voice, speech recognition, speech generation, audio generation, and music.
The headline numbers are self-reported in the paper and mirrored on Artificial Analysis's public board. No independent replication has been published yet.
Originally reported by paper
Read the original article →Original headline: StepFun's StepAudio 3 Realtime Claims #1 on Artificial Analysis Full-Duplex Bench