Alibaba Ships Qwen-Audio 3.1 Stack, Cuts Voice APIs Up to 95%
TL;DR
- Qwen-Audio ASR fell 95%, TTS 70%, and Realtime 85%, applying the same token commoditization playbook Alibaba ran against OpenAI on LLM pricing.
- TTS-Next unifies language modeling and diffusion in a single pass, eliminating the three-API orchestration chain traditional voice pipelines require.
- At 20M characters per month, Qwen TTS costs low hundreds of dollars versus roughly $2,000 at ElevenLabs, a gap wide enough to force procurement reviews.
Alibaba's Qwen team pushed a new voice stack and, in the same post, cut prices on its audio APIs by up to 95 percent. The-Decoder reports that ASR drops by up to 95 percent, TTS by about 70 percent, and Realtime by roughly 85 percent.
The lineup is five models: upgraded versions of ASR, TTS and Realtime, plus two additions, TTS-Next for audio creation and ASR-Next for audio understanding.
The capability descriptions are specific. The upgraded ASR "automatically cleans up filler words and repetitions" and improves multilingual and dialect recognition. ASR-Next layers in multi-speaker identification with timestamps, emotion detection, and ambient sound and machine noise detection. TTS delivers "multilingual synthesis with natural cross-language voice transfer," with emotion, speed and style set through text prompts. TTS-Next generates "voice, sound effects, and background audio in a single pass." The Realtime model handles simultaneous speaking and listening with "instant interruption."
The drop lands in a busy stretch of Alibaba AI coverage, and slots alongside the other new voice releases moving through our Voice AI feed this week from Google and OpenAI. The article does not include an effective date for the discounted pricing or say whether the tier is capacity-committed or spot.
What others are reporting
-
OrcaRouter Read →
Granular price modeling: 20M chars/month costs low hundreds at Qwen vs ~$2,000 at ElevenLabs; flags token-based realtime billing as an unpredictable savings wildcard.
A 70% cut against a $100-per-million-character incumbent is not a rounding difference.
-
AlphaSignal Read →
Technical breakdown of TTS-Next's unified LM+diffusion architecture; notes it collapses three API handoffs into one and benchmarks prior model at 1,237 Elo vs ElevenLabs.
Fewer orchestration boundaries can reduce latency and simplify interruption handling, especially when a user speaks over the agent.
-
Data Studios Read →
Frames the 262K context window as an agent runtime shift: native function calling and web search make Realtime Plus a stateful voice agent, not a voice I/O wrapper.
Increasing available context gives developers substantially more room before older information must be removed or summarized.
-
tbreak Read →
Middle East-focused lens: Alibaba has not published UAE-specific pricing or data-handling terms, flagging a regional deployment gap for enterprise buyers outside China.
Lower per-call costs reduce the pressure to limit usage for high-volume transcription and customer service applications.
-
PANews Read →
Positions the cuts as a competitive accessibility move for Asian-market developers and enterprises, foregrounding cost reduction over technical specifications.
Prices across the entire Qwen-Audio voice model lineup have been lowered, with TTS cut ~70%, Realtime ~85%, and ASR as much as 95%.
Originally reported by the-decoder.com
Read the original article →Original headline: Alibaba Ships Qwen-Audio 3.1 Stack With TTS-Next and ASR-Next, Cuts Voice API Prices Up to 95%