Alibaba’s Qwen team released Qwen-Audio-3.1, a five-model voice stack, and Alibaba Cloud cut the rates of its realtime and text-to-speech (TTS) models from September 22, 00:00 Beijing time. This is a Desk Bot briefing from Alibaba Cloud’s price notice, its Model Studio documentation and the Qwen team’s own announcement.
The realtime price cut, from Alibaba’s notice
Alibaba Cloud’s price reduction notice, published September 21, lists old and new rates for its International (Singapore) site. The realtime model it names is qwen-audio-3.0-realtime-flash; the notice does not list a 3.1 realtime price.
| qwen-audio-3.0-realtime-flash | Before (USD / 1M tokens) | After (USD / 1M tokens) | Cut (desk’s arithmetic) |
|---|---|---|---|
| Input text | 0.45 | 0.23 | 49% |
| Input audio | 4.50 | 0.93 | 79% |
| Output text | 4.50 | 0.70 | 84% |
| Output text + audio | 15.00 | 1.87 | 88% |
The audio lines, which dominate a voice call’s bill, fall by 79% to 88%. That fits the Qwen team’s “Realtime ~85% off” in its launch post. The notice says no action is needed and that usage before the effective time keeps the old prices.
TTS and ASR: the percentages are Qwen’s
For qwen-audio-3.1-tts-flash the notice changes the billing unit rather than just the rate: from $0.15 per 10,000 input characters to $0.23 per million input tokens plus $1.87 per million output tokens. A character-based price and a token-based price cannot be compared without knowing how many tokens your text and audio produce, so the desk cannot confirm the “TTS ~70% off” figure. It stays Qwen’s claim. On the China (Beijing) region, the model page lists 1.5 CNY input and 12 CNY output per million tokens.
“ASR up to 95% off” appears only in Qwen’s post. The price notice does not cover speech recognition. The qwen-audio-3.1-asr-flash page shows current rates, but no old price to compare them with:
| qwen-audio-3.1-asr-flash | Input (USD / 1M tokens) | Output (USD / 1M tokens) |
|---|---|---|
| Singapore | 0.15 | 0.47 |
| China (Beijing) | 0.113 | 0.382 |
What the five models are
Alibaba Cloud’s Apsara Conference press release of September 22 names the upgraded ASR, TTS and Realtime models and the new TTS-Next. Qwen’s post adds ASR-Next. By Qwen’s own description:
- ASR handles more languages and dialects and removes filler words from transcripts.
- ASR-Next labels speakers with timestamps and picks up emotions, ambient sounds and machine noise.
- TTS takes plain-language instructions for emotion, speed and style.
- TTS-Next generates speech, sound effects and background audio in one pass. Its model page lists Beijing-only pricing of $0.848 input and $1.696 output per million tokens, Chinese and English only.
- Realtime listens while it speaks and can be interrupted. The realtime guide lists qwen-audio-3.1-realtime-plus.
The desk found no Model Studio page for ASR-Next yet. None of these quality claims has been independently tested.
Who should care
Teams running voice agents or call bots on Alibaba Cloud should check their bill now: the realtime cut is large and already applies. Anyone comparing TTS or ASR providers should run their own text and audio through the token counter first, since the headline percentages cannot be checked against a like-for-like old price.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.