Alibaba Ships Qwen-Audio 3.1 Stack, Cuts Voice APIs Up to 95%
TL;DR
- Alibaba's Qwen team released Qwen-Audio-3.1, a five-model lineup covering ASR, TTS, Realtime, plus new TTS-Next for creation and ASR-Next for understanding.
- Speech recognition API pricing drops by up to 95 percent, TTS by about 70 percent, and the Realtime voice model by roughly 85 percent.
- ASR-Next adds multi-speaker identification with timestamps, emotion detection, and ambient sound recognition; TTS-Next generates voice, effects, and background audio in a single pass.
Alibaba's Qwen team pushed a new voice stack and, in the same post, cut prices on its audio APIs by up to 95 percent. The-Decoder reports that ASR drops by up to 95 percent, TTS by about 70 percent, and Realtime by roughly 85 percent.
The lineup is five models: upgraded versions of ASR, TTS and Realtime, plus two additions, TTS-Next for audio creation and ASR-Next for audio understanding.
The capability descriptions are specific. The upgraded ASR "automatically cleans up filler words and repetitions" and improves multilingual and dialect recognition. ASR-Next layers in multi-speaker identification with timestamps, emotion detection, and ambient sound and machine noise detection. TTS delivers "multilingual synthesis with natural cross-language voice transfer," with emotion, speed and style set through text prompts. TTS-Next generates "voice, sound effects, and background audio in a single pass." The Realtime model handles simultaneous speaking and listening with "instant interruption."
The drop lands in a busy stretch of Alibaba AI coverage, and slots alongside the other new voice releases moving through our Voice AI feed this week from Google and OpenAI. The article does not include an effective date for the discounted pricing or say whether the tier is capacity-committed or spot.
Originally reported by the-decoder.com
Read the original article →Original headline: Alibaba Ships Qwen-Audio 3.1 Stack With TTS-Next and ASR-Next, Cuts Voice API Prices Up to 95%