← All posts / Models

Five Models and a 95% Price Cut: Alibaba's Qwen-Audio-3.1 Resets the Voice AI Cost Floor

Alibaba's Qwen team ships a five-model voice stack — upgraded ASR, TTS and Realtime plus new TTS-Next and ASR-Next — while slashing ASR prices by up to 95%, TTS by ~70% and Realtime by ~85%.

Five Models and a 95% Price Cut: Alibaba's Qwen-Audio-3.1 Resets the Voice AI Cost Floor

Alibaba’s Qwen team has never been shy about competing on price, but its latest release still lands like a shockwave. On September 23, the team quietly pushed out Qwen-Audio-3.1, a coordinated refresh of its entire voice stack: five models covering speech recognition, speech synthesis, and live conversation, accompanied by API price cuts of up to 95 percent. In a market where OpenAI, ElevenLabs, and Deepgram charge premium rates for audio intelligence, Alibaba has just reset the cost baseline the entire transcription and voice category must now meet.

What shipped

The release is not one model but a coordinated lineup. Three existing models — ASR (automatic speech recognition), TTS (text-to-speech), and Realtime — received full upgrades, and two new models joined the family:

  • ASR (upgraded) — The workhorse transcription model now “automatically cleans up filler words and repetitions,” producing cleaner transcripts without post-processing, alongside improved multilingual and dialect recognition.
  • ASR-Next (new) — The understanding tier. It layers multi-speaker identification with per-speaker timestamps, emotion detection, and ambient-sound and machine-noise detection on top of transcription. In other words: not just what was said, but who said it, when, in what emotional state, and against what background.
  • TTS (upgraded) — Multilingual synthesis with “natural cross-language voice transfer,” with emotion, speed, and style controllable through plain text prompts rather than opaque parameter knobs.
  • TTS-Next (new) — The most architecturally interesting model in the stack. Described in Alibaba Cloud’s documentation as an “AudioGen” model, it generates voice, sound effects, and background audio in a single pass, collapsing what previously required three separately orchestrated inference systems into one model call.
  • Realtime (upgraded) — Handles simultaneous speaking and listening with “instant interruption” handling, the full-duplex capability that separates conversational voice agents from walkie-talkie-style assistants.

The technical specifics from Alibaba Cloud Model Studio flesh out TTS-Next’s ambitions. It accepts text, timestamps, and up to three reference audio clips (30 seconds each), outputs WAV/MP3/PCM audio, and supports tasks spanning multilingual TTS, multi-speaker dialogue, podcasts, cinematic soundscapes, ambient audio, and sound effects. Output duration per request reaches 240 seconds for podcasts, 120 for other scenarios. List pricing runs 6 CNY per million input tokens and 12 CNY per million output tokens before promotions.

The price cut is the story

Model releases are routine; this pricing is not. Per The Decoder’s report, the cuts break down as:

APIPrice cut
ASRup to 95%
TTS~70%
Realtime~85%

A 95 percent cut is not a promotion — it is a strategic repricing of an entire category. AI Weekly’s analysis put it bluntly: the cut “resets the cost baseline the entire transcription vendor category must meet, regardless of what prior margins were built on.” Western vendors like Deepgram, ElevenLabs, and OpenAI’s audio APIs now face an uncomfortable question: match the new floor and destroy their margin structure, or hold prices and concede every price-sensitive workload — call centers, meeting transcription, media monitoring, voice agents at scale — to Alibaba.

There are caveats worth noting honestly. Alibaba did not publish an effective date for the discounted pricing, nor whether the tier is capacity-committed or spot. The aggressive list prices also apply to Model Studio, Alibaba Cloud’s own inference platform, meaning the full discount lands for customers willing to build on Chinese cloud infrastructure — a real consideration for regulated Western enterprises. But for the global majority of developers, the signal is unambiguous: voice AI is heading toward effectively-free transcription.

Why TTS-Next matters architecturally

Beyond pricing, TTS-Next points at where audio generation is going. Today, producing a polished podcast or video soundtrack means chaining an orchestration pipeline: one model for narration, a second for sound effects, a third for ambient beds, then manual mixing. TTS-Next’s single-pass generation of speech, effects, and background audio collapses that stack — the same architectural consolidation that multimodal LLMs brought to text-and-image workflows.

The “flexible timing control” via timestamp inputs is similarly practical: creators can pin narration to specific moments without iterative re-generation. Combined with consistent voice identity across long outputs and multi-speaker dialogue support, the model targets professional production scenarios — short-form video, podcasts, film and game audio — rather than demo-ware.

Meanwhile, the upgraded Realtime Plus carries a 262K-token context window, enough to hold conversation history, tool results, and application state across a full session. That extends Alibaba’s competitive surface beyond transcription vendors into real-time agent orchestration — territory currently contested by OpenAI’s Realtime API and a swarm of voice-agent startups.

Context: the voice race is heating up

Qwen-Audio-3.1 does not land in a vacuum. The same week saw Nvidia release Nemotron 3 Diarization, a free 100M-parameter model that identifies up to eight speakers in real time and currently leads the VoiceArena diarization benchmark at a 14.7% error rate. OpenAI shipped three ChatGPT Voice upgrades, including plugin support and user-selectable backends. Google, per AI Weekly’s tracking, has its own voice releases moving through the pipeline.

What differentiates Alibaba’s move is the combination: frontier-adjacent capabilities and an aggressive price umbrella. Nvidia’s diarization model is free but narrow. OpenAI’s voice upgrades are polished but premium-priced. Qwen-Audio-3.1 is betting that in voice — more than in text — price elasticity dominates, and that the explosion of agent workloads will flow to whoever makes audio intelligence nearly free.

For enterprises building voice agents, meeting bots, or media pipelines, the practical advice is straightforward: benchmark Qwen-Audio-3.1 against your current stack now. Even if data-residency or compliance rules keep you on a Western vendor, the new price floor belongs in every renegotiation. For the rest of the industry, the release is another reminder that in the current AI cycle, the most disruptive product feature is frequently the invoice.