← All posts / Models

Qwen3.8-Flash-Next: Alibaba Open-Sources the First Glimpse of Qwen4's Architecture

Alibaba's Qwen team open-weights a 125B-parameter MoE that activates just 6B per token, pairs 1M-token context with a hybrid GDN + QSA attention stack, and beats Claude Opus 4.6 Max on agentic coding benchmarks.

Qwen3.8-Flash-Next: Alibaba Open-Sources the First Glimpse of Qwen4's Architecture

On August 26, 2026, Alibaba’s Qwen team quietly dropped one of the more consequential open-weight releases of the year: Qwen3.8-Flash-Next, a 125-billion-parameter multimodal mixture-of-experts model that activates only about 6 billion parameters per token. The “Next” suffix is doing real work here — this is explicitly framed as an early preview of the architecture that will underpin Qwen4, making it the first public artifact of Alibaba’s next-generation model design. For anyone tracking where frontier-adjacent open models are heading, this release is less a product launch and more a research manifesto about cost-efficiency.

What makes the architecture new

Qwen3.8-Flash-Next upgrades the model systematically along four axes — attention, residual connections, embedding, and optimization — with the attention redesign being the headline. The team built a hybrid attention stack that mixes two very different mechanisms across the model’s 48 layers:

  • Gated DeltaNet (GDN) in 36 of the 48 layers. GDN is a linear-attention variant whose key/value state stays roughly constant as context grows, which is what makes million-token inference economically plausible. The three-to-one ratio of GDN to full-attention layers is the structural bet: the vast majority of tokens flow through the cheap path.
  • QSA (queried sub-attention, per the team’s description) in the remaining 12 layers. Unlike the GDN layers, QSA layers’ K/V state does grow with context — but QSA operates at the micro-block level rather than over the full sequence, using sparse queries over token blocks (with linear attention heads numbering 48 for value and 16 for query), which cuts long-context latency significantly without abandoning precise retrieval entirely.

The design philosophy is pragmatic: pure linear attention degrades on tasks that need exact long-range recall, while pure full attention becomes prohibitively expensive at 1M tokens. Qwen’s answer is to make full-precision attention a scarce resource — 12 layers’ worth — and spend the rest of the compute budget on a constant-memory approximation. Notably, Qwen3-Coder-Next took a similar hybrid approach earlier this year, but Flash-Next pushes the sparsity and context ambitions considerably further.

The model ships with 262K native context, extensible to 1M tokens via YaRN, and it is fully multimodal — text and vision in one set of weights, no separate vision adapter tier.

Benchmarks: small active footprint, agentic coding punch

The most eye-catching numbers are on agentic coding, where the model is clearly optimized. Qwen’s published results include:

  • 62.5 on SWE-bench Pro — ahead of Claude Opus 4.6 Max at 53.4 on the same harness
  • 58.7 on DeepSWE 1.1, a sustained-horizon software-engineering eval
  • 73+ on CoWorkBench, the office-workflow agent benchmark family
  • 81.0 on SWE-bench (verified)

Caveats worth stating plainly: these are Qwen’s own harness numbers, and the team notes that problematic tasks were corrected and all baseline models re-evaluated on the refined benchmark — which is disclosed, but still means comparisons rest on Qwen’s re-runs rather than a neutral third party. The pattern, however, is consistent across every eval family that was published: a model with a 6B-active-parameter footprint is landing in territory previously occupied by far denser frontier systems, specifically in agentic and coding tasks rather than pure knowledge recall.

The economics story

The release blog post is titled “A New Architecture, Towards Ultimate Cost-Efficiency,” and the pricing backs the rhetoric. On QwenCloud, the production API version — qwen3.8-flash — is priced at roughly $0.16 per million input tokens and $0.47 per million output tokens. Reuters reported that the model delivers stronger coding and office-task performance while cutting training costs substantially relative to the previous Flash generation. When a 6B-active MoE hits frontier-adjacent agentic benchmarks, the cost per solved task drops by an order of magnitude, which is precisely the kind of shift that moves workloads.

One nuance the community flagged quickly: Qwen3.8-Flash-Next is the open-weight release, while Qwen3.8-Flash is the production API version with the 1M context window and built-in tool support. The open weights on Hugging Face print a slightly different parameter accounting than the API tier (the HF model card lists ~180B total for the base configuration in some views, with 125B/A6B as the commonly cited post-training figure), so deployers should check the specific checkpoint they pull.

Availability

The weights are live under Apache 2.0 on Hugging Face (Qwen/Qwen3.8-Flash-Next) and already on Ollama — including MLX builds for Apple Silicon (the 125b-a6b-mlx-bf16 tag), meaning a sufficiently provisioned Mac Studio can run a Qwen4-architecture preview locally. Ollama’s library page describes it as “the first open-weight model built on the architecture that will underpin Qwen4,” which matches Qwen’s own framing. Community reception on r/LocalLLaMA has been substantial, with release-day megathread users testing it on single-GPU setups like DGX Spark.

Why this matters

Three takeaways for anyone watching the open-model ecosystem:

  1. Architecture is the new scaling lever. With training compute costs under pressure, Qwen is competing on inference economics rather than raw parameter counts. A 3:1 GDN-to-QSA hybrid ratio at 1M context is a concrete, copyable design point, and Apache 2.0 means everyone from startups to rival labs can study it.
  2. The agentic coding frontier is now contested by open weights. Beating Claude Opus 4.6 Max on SWE-bench Pro — even on a self-run harness — was unthinkable for an open model twelve months ago. It signals that Anthropic’s strongest commercial moat is under genuine pressure on the margin.
  3. Qwen4 now has a public shape. By open-sourcing the architecture preview, Alibaba is both stealing mindshare ahead of its next flagship launch and effectively crowdsourcing validation: every community deployment of Flash-Next is a stress test of the Qwen4 foundation.

The counterpoint is durability. Hybrid linear-attention models have historically shown quality cliffs on certain long-context recall tasks, and early single-device reports mention degraded behavior in some query regimes. Whether the 12 QSA layers are enough to prevent those cliffs at 1M tokens is exactly what the next few weeks of community evaluation will determine — and, conveniently for Alibaba, that evaluation is happening on their future flagship’s architecture, for free.