Shanghai Jiao Tong Open-Sources AuK, a 1.5B Foundation Model Unifying Speech Generation and Editing
Summary
The AuK technical report describes a 1.5B open-source foundational model trained on ~3B instruction-audio pairs and 1.95M hours of audio to handle zero-shot TTS, acoustic and paralinguistic editing, content edits, enhancement and separation via one natural-language interface. AuK posts 2.65% WER on Seed-TTS-Eval (vs 3.07% for Qwen3-TTS) and cuts Chinese speech-editing WER from 10.46% to 3.09%. Code and weights are released.
Originally reported by huggingface.co
Read the original article →Original headline: Shanghai Jiao Tong Open-Sources AuK, a 1.5B Foundation Model Unifying Speech Generation and Editing