paper web signal

StepAudio 3 Music drafts ABC notation before 5:30 tracks

TL;DR

  • StepAudio 3 Music first produces an ABC-notation arrangement plan the paper calls ABC-CoT, then predicts audio tokens decoded to 48-kHz sound.
  • The system supports song, instrumental, accompaniment-from-dry-vocals and cover-song generation for up to 5 minutes and 30 seconds.
  • On the Artificial Analysis Music Arena Vocals leaderboard the authors report a Quality Elo of 1105, behind Suno V5.5 and Mureka.

The system writes sheet music before it writes audio. In a new arXiv technical report, the authors of StepAudio 3 Music describe a Mixture-of-Experts autoregressive model that uses "ABC notation to produce an intermediate arrangement plan (ABC-CoT) before predicting music tokens," so "harmony, rhythm, and melodic structure" enter the generation context as a readable score rather than as opaque latents. Only then does a flow-matching diffusion Transformer predict continuous VAE latents that a decoder turns into 48-kHz audio.

The stated ceiling is five and a half minutes. The abstract says the training curriculum and supervised fine-tuning "support song and instrumental generation, accompaniment generation from dry vocals, and cover-song synthesis for up to 5 minutes and 30 seconds," with a further round of reinforcement learning via direct preference optimization on top.

On ranked evaluation the authors are specific about where they land. They report "a Quality Elo of 1105, behind only Suno V5.5 and Mureka and ahead of Suno V5, MiniMax models, and other systems" on the preliminary Artificial Analysis Music Arena Vocals leaderboard. The paper also claims "the highest AudioBox Content Enjoyment, Content Usefulness, and Production Quality scores" among the systems it evaluates, alongside the highest MuQ-MuLan similarity and what it describes as "competitive SongBench results."

The abstract publishes no per-benchmark numbers, does not enumerate the systems it was compared against beyond the leaderboard entries, and does not describe the training corpus or any licensing arrangement around cover-song synthesis.