95% Off the Sound of AI: Alibaba's Qwen-Audio 3.1 Declares a Voice Price War
Alibaba's Qwen team ships a five-model audio stack — ASR, ASR-Next, TTS, TTS-Next, and a 262K-context Realtime model — while cutting API prices by up to 95%, collapsing the cost of voice agents overnight.
The loudest AI news this week isn’t a frontier chatbot — it’s the sound of a price war. Alibaba’s Qwen team has released Qwen-Audio 3.1, a coordinated five-model lineup covering speech recognition, speech synthesis, and real-time conversation, and alongside it came cuts to API pricing that are hard to read as anything other than an attempt to commoditize voice AI: text-to-speech drops roughly 70%, real-time conversation roughly 85%, and automatic speech recognition by up to 95%.
For an industry racing to put AI agents on the phone, in earbuds, and inside customer-support queues, audio was quietly becoming a major line item. Alibaba just made it nearly free — and in doing so, told us something important about where the voice-agent market is heading.
What shipped: five models, one stack
Qwen-Audio 3.1 isn’t a single model but a coordinated lineup that covers the entire voice loop:
- ASR — the workhorse transcription model. It handles multilingual and dialect recognition and adds a “polishing” pass that automatically strips filler words and repetitions from transcripts.
- ASR-Next — the audio understanding model, and arguably the most interesting of the five. It performs speaker diarization (assigning segments to individual speakers), generates timestamps, detects emotions, recognizes ambient and machine sounds, supports sound captioning, and can answer questions about audio content directly.
- TTS — the speech generation model, with multilingual synthesis, cross-lingual voice transfer, and instruction-based control over emotion, speed, accent, and delivery style. You can direct it with plain text: “Read this with a sharp, commanding tone, demanding respect.”
- TTS-Next — a unified audio-generation system that pairs a language model with a diffusion approach to produce speech, sound effects, and background audio in a single pass. Alibaba’s documentation describes it as an “AudioGen” model aimed at podcasts, cinematic soundscapes, and game audio rather than simple narration.
- Realtime — the conversational model, which speaks and listens simultaneously (full duplex), handles interruptions instantly, and — in a detail that says a lot about where voice UX is going — responds more slowly and with more empathy when it detects a low mood in the user’s voice.
The flagship Realtime Plus variant expands the context window to 262,144 tokens and supports function calling, web search, and voice cloning over persistent realtime connections, streaming speech and text responses concurrently.
The technical story: fewer tokens, faster training
The accompanying technical paper gives the clearest picture of how the TTS model gets its economics. A low-frame-rate tokenizer represents speech at just 12.5 frames per second, sharply reducing the number of units the autoregressive language model has to predict. A separate flow-matching component then generates the acoustic detail needed for the final waveform.
Training runs in five stages: pretrain the language model, pretrain the flow-matching model, train both jointly while progressively concentrating on higher-quality data, then apply reinforcement learning to each component in turn. Alibaba says the RL stages specifically improve prosody, voice similarity, and resilience to difficult reference recordings — the qualities that make synthesized speech feel human rather than merely intelligible.
The model accepts free-form instructions for role, emotion, speaking rate, timbre, style, and accent, and supports 86 inline tags for nonverbal events like laughter, breathing, coughing, and sighing. Language coverage spans 16 languages and 20 Chinese dialect regions, with single-pass synthesis of up to three minutes of audio (four minutes for podcast scenarios on TTS-Next, per Alibaba Cloud’s documentation) and tolerance for noisy, reverberant, or unclear reference speech.
The pricing story: read it next to DeepSeek
The price cuts are the headline, but they make the most sense placed beside another number from this week: DeepSeek hitting $1 billion in annual recurring revenue after raising API prices 2.3x to 4.5x. The two moves are opposite in direction and identical in meaning. The Chinese AI ecosystem has bifurcated into labs that can charge a premium for frontier reasoning, and labs that treat audio — and other commodity modalities — as volume businesses where the cheapest capable provider wins.
There is historical context for Qwen’s confidence here. In July 2026, the previous Qwen-Audio 3.0 TTS Plus ranked first in the Artificial Analysis Text-to-Speech arena with a Quality Elo score of 1,237, ahead of Google’s Gemini 3.1 Flash TTS, MiniMax Speech 2.8 HD, and ElevenLabs Eleven v3 — while listing at $27.60 per million characters, roughly one-third of the compared ElevenLabs and MiniMax tiers. (Caveats apply: that model generated about 16 characters per second versus 120 for the fastest competitor, and its lead over the runner-up fell within overlapping confidence intervals — a statistical tie at the top.)
The new endpoints still need independent quality and latency measurement. But if the 3.1 line holds anywhere near that quality-per-dollar position at the new prices, the margin structure of every independent voice-AI vendor just got stress-tested.
Why full duplex matters more than benchmarks
The most consequential model in the lineup may be the least glamorous: Realtime. Conventional voice agents chain together a pipeline — speech → ASR → text → language model → text → TTS → speech — and every handoff adds latency and failure surface, especially when a user talks over the agent mid-sentence. The integrated model folds speech understanding, reasoning, and generation into one service, which simplifies interruption handling and cuts end-to-end delay.
Alibaba notes that applications built on Qwen-Audio 3.0 Realtime Plus can keep the same integration protocol, limiting migration code. Production teams will still need regression tests for turn-taking, tool-call accuracy during interruptions, and error handling — and the 262K context window is a session ceiling, not a memory strategy; live audio, tool results, and conversation state all consume it, so stateful agents still need summarization and truncation policies.
What’s open, what’s hosted, what’s missing
A detail many coverage pieces skipped: availability is uneven across the lineup. The ASR file-transcription endpoint and Realtime Plus are live hosted APIs on Alibaba Cloud Model Studio today. The open-weight Qwen3-TTS repository remains available under Apache 2.0 — but it belongs to a separate lineage from the hosted Qwen-Audio 3.x services, so developers eyeing self-hosting should verify feature and API parity before assuming equivalence. ASR-Next and TTS-Next have no launch dates yet; they’re announced but not generally available. TTS-Next pricing is published (6 CNY per million input tokens, 12 CNY per million output tokens in Beijing, excluding promotions), with a 3 RPS rate limit that suggests early-capacity caution.
The takeaway
Qwen-Audio 3.1 is a quiet kind of significant. No single model here rewrites the frontier — but a five-model stack with arena-leading TTS quality, full-duplex realtime conversation with a 262K context, speaker-aware audio understanding, and price cuts of 70–95% changes the arithmetic for everyone building voice products. When the marginal cost of a transcribed minute or a synthesized paragraph approaches zero, the differentiators move elsewhere: latency, reliability, tool-use under interruption, and the taste to direct a voice that people actually want to hear.
The capability news this month has been loud. The pricing news says more: voice is becoming infrastructure, and infrastructure gets cheap fast.
Sources
- [1] https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/
- [2] https://alphasignal.ai/news/alibaba-s-qwen-audio-3-1-slashes-voice-api-prices-by-up-to-95
- [3] https://x.com/Alibaba_Qwen/status/2102687258990026993
- [4] https://help.aliyun.com/en/model-studio/qwen-audio-3-1-tts-next