Alibaba's Qwen cuts audio API prices, ASR down as much as 95%
Voice API prices keep falling, with ASR cut the deepest; all capability claims are Qwen's own, so cost assumptions can be repriced now.
Original event 2026-09-23
Alibaba's Qwen team released Qwen-Audio-3.1, a lineup of five audio models, and cut its voice API prices: TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent.
The series covers speech recognition, text-to-speech and real-time interaction. Qwen says the ASR model improves multilingual and dialect recognition and cleans up filler words, TTS-Next generates voice and sound effects in a single diffusion pass, and the real-time model supports simultaneous speaking and listening with instant interruption.
All capability claims are Qwen's own; per The Decoder, they come from the official blog and an X announcement, with no independent evaluation yet. Cost assumptions for voice applications can be repriced at the new rates.