One Model to Hear Everything: Qwen's Qwen3.8-Omni-Flash Cuts Audio Costs 98% and Reads 2-Hour Video Like an Agent
Alibaba's Qwen team ships a native omnimodal model with a 1M-token context, 74-language speech recognition, and a 98% cut in per-hour audio input cost — plus an agentic video mode that skips 45% of the frames and still scores higher.
On September 18, 2026, Alibaba’s Qwen team released Qwen3.8-Omni-Flash, the newest member of its omni family and arguably its most aggressive attempt yet to make audio and video first-class citizens of the agent economy. The pitch is simple to state and hard to execute: one natively end-to-end model that takes text, images, audio, and video in a single request, holds a 1 million-token context window, and costs so little per hour of listening that the separate speech-to-text pipeline stops making economic sense.
The timing matters. For two years, “multimodal” has mostly meant “a language model that can also look at pictures.” Audio and video — the modalities that actually dominate meetings, calls, media, and the physical world — stayed expensive, lossy, and architecturally second-class: a transcription service feeding text into a chatbot, stripping tone, overlapping speakers, and everything on screen. Qwen3.8-Omni-Flash is the latest signal that this era is ending, and that the battleground is shifting from perception to agentic delivery: models that don’t just describe a two-hour recording, but decide which parts of it to act on.
What shipped
The headline numbers, all from Qwen’s own release post:
- Native omnimodal architecture — text, image, audio, and video input with text output, in one end-to-end model rather than a cascade of specialist components. Maximum output is 131,072 tokens.
- 1M-token context — enough to hold roughly two hours of raw audio or video per file, with up to 64 files per request on Alibaba Cloud Model Studio.
- 74 languages for speech recognition, from Arabic and Cantonese to Uyghur and Vietnamese, and 29 languages for speech generation.
- +25% average score across 29 benchmarks versus the previous-generation Qwen3.5-Omni-Plus, spanning audio, audio-visual, and agent evaluations.
- Over 98% lower API cost per hour of audio input, and over 93% lower for combined audio-visual input, compared with Qwen3.5-Omni-Plus using Qwen’s own estimation method.
That last claim deserves careful reading, and Qwen deserves credit for publishing the method. The 98% figure is a like-for-like comparison against its own previous omni model, estimated as 30× the input cost of a two-minute clip — not a comparison with OpenAI or Google. But working from Alibaba Cloud’s published pricing makes the magnitude plain: Model Studio bills audio at 7 tokens per second and lists the model internationally at USD 0.15 per million input tokens. One hour of audio is 25,200 tokens. That works out to roughly $0.0038 per hour of audio input — well under a cent, before output costs. Audio-visual input at 720p and 1 fps is down more than 93% on the same basis.
The agentic video mode
The most technically interesting idea in the release isn’t a benchmark score — it’s a search policy. When a long recording is analyzed the usual way, the entire file is processed even when the answer sits in three minutes of it. Qwen3.8-Omni-Flash’s agentic mode inverts this: the model starts from the question, decides which segments to watch and listen to, and gathers evidence in several passes from coarse to fine, so most frames are never processed at all.
The numbers Qwen reports are striking. On OmniVideoBench, accuracy rose from 63.4 in the static setting to 67.8 in the agent setting while tokens per query fell from 145,736 to 79,117 — a reduction of about 45.7%. Fewer tokens, better answers. The caveat is that the agent setting ran inside Qwen Code, Qwen’s own coding-agent harness, and Gemini 3.8 Flash — put through the same harness — scored 70.1 in agent mode, ahead of Qwen’s 67.8, while Qwen led on the long-video LVOmniBench at 73.6 against 70.7. The pattern, though, transfers to any long-media product: let the model index first and read selectively, and the bill falls while accuracy holds.
The benchmark picture, honestly read
Qwen’s comparison table mixes two opponents — its own previous model, where the deltas are dramatic, and Gemini 3.8 Flash, where the picture is genuinely mixed:
| Benchmark | Qwen3.8-Omni-Flash | Qwen3.5-Omni-Plus | Gemini 3.8 Flash |
|---|---|---|---|
| WildClawBench-MM (multimodal tool use) | 71.0 | 34.5 | 58.9 |
| AgenticVBench (multimodal tool use) | 36.8 | 14.5 | 45.0 |
| UniClawBench (multimodal tool use) | 69.6 | 67.1 | 69.0 |
| OmniVideoBench (audio-visual reasoning) | 63.4 | 53.8 | 65.2 |
| LongAudioSpan (long-audio accuracy) | 82.7 | 74.4 | 79.3 |
| AliMeeting (speaker error / word error, lower is better) | 3.4 / 17.2 | 88.1 / 89.6 | 72.6 / 53.1 |
| FLEURS-ASR (60-language word error, lower is better) | 9.3 | 7.2 | 7.9 |
| VoiceBench (audio interaction) | 91.6 | 92.9 | 92.3 |
The standout row is multi-speaker meetings. Qwen reports its previous model essentially failing the AliMeeting transcription test (speaker error 88.1) and the new one bringing that to 3.4 with word error at 17.2 — the single capability behind the meeting-minutes use case. Also notable: the two rows where the new model trails its own predecessor (multilingual transcription at 9.3 versus 7.2, and VoiceBench at 91.6 versus 92.9) are printed rather than hidden. Two of the agent benchmarks were run inside the Claude Code harness and one inside OpenClaw, so those scores measure model-plus-harness — a distinction that can move results by a wide margin.
Qwen’s own summary is fair: close to Gemini 3.8 Flash on audio-visual work, ahead of it on audio overall. On AliMeeting it’s not close at all — it’s a generational gap.
The demo that explains the philosophy
Buried in the release post is an anecdote that does more explanatory work than any chart: Qwen gave its own model a task — improve Qwen2.5-Omni-3B’s Sichuan-dialect speech recognition within 12 hours and deliver a usable model. The omni model wasn’t the subject of the experiment; it was the researcher. It read the codebase, planned a data strategy for a regional Chinese dialect, ran training, and shipped — audio understanding as an input to autonomous ML engineering, not just an output of it.
That framing — “Omni Senses. Agentic Delivery.” — is the actual thesis of the release. Audio and video are evolving from perceptual inputs into core media that agents use to understand their environment and execute tasks. To back it, Qwen also expanded Qwen-MM-Plugins, its open-source skill-and-tool-server set that lets coding agents like Claude Code, Codex, and Qwen Code read images, video, and audio, and open-sourced the Qwen-Live Harness, a real-time omnimodal runtime built on the Flash-Realtime API (at time of writing, its GitHub page still returns a 404).
The catch: no open weights
For a lab that built its reputation on open weights, the quiet part of this release is what it doesn’t say. There is no open-weights announcement and no model card on Hugging Face as of September 18. Until Qwen says otherwise, Qwen3.8-Omni-Flash is an API-only model, served through Alibaba Cloud Model Studio and the Qianwen platform. It is also distinct from Qwen3.8-Max, the text-and-vision flagship released earlier this month — the Omni model is the one that hears.
That positioning says something about where the frontier of competition now sits. The cheap, commodity part of multimodality — transcribe, then chat — is being commoditized to fractions of a cent per hour. The differentiated part is agentic: models that search long media selectively, call tools mid-stream, and hold a two-hour recording in context while they work. Qwen3.8-Omni-Flash is a credible claim on that frontier, at a price that makes building for it hard to resist.
What to watch
For teams building meeting-minutes, call-review, captioning, dubbing, or long-video search products, the decision this release changes is whether to keep a separate speech-to-text step at all. The sensible test is unglamorous: run the same ten recordings through your current pipeline and through this model, and score them on cost per hour and error count. But the wider race is worth watching too. If Gemini 3.8 Flash’s agent-mode lead on OmniVideoBench holds under independent evaluation, Google retains the crown on audio-visual reasoning; if Qwen’s meeting-transcription gap and price cuts force a response, the omni-agent market — where the marginal cost of listening approaches zero — opens up for everyone.
Sources
- [1] https://qwen.ai/blog?id=qwen3.8-omni-flash
- [2] https://gigazine.net/gsc_news/en/20260918-qwen-3-8-omni-flash/
- [3] https://www.digitalapplied.com/blog/qwen3-8-omni-flash-omnimodal-agents-audio-video-cost
- [4] https://www.alibabacloud.com/blog/qwen3-omni-flash-2025-12-01-hear-you--see-you--follow-smarter_602735