← All posts / Tools

vLLM 0.28.0 Lands Decode Context Parallel, DFlash2 Speculative Decoding and Tiered Disk KV Offload

The vLLM project's latest release packs 584 commits from 270 contributors: Decode Context Parallel for Kimi-K3, end-to-end sparse MLA for DeepSeek V4, DFlash2 speculation, disk-tier KV offloading, and doubled batch defaults.

vLLM 0.28.0 Lands Decode Context Parallel, DFlash2 Speculative Decoding and Tiered Disk KV Offload

The quiet workhorse of open-source AI serving just shipped its biggest update of the summer. vLLM v0.28.0, published August 26, 2026, merges 584 commits from 270 contributors — 76 of them first-time contributors — into a release that reads less like a routine version bump and more like a coordinated performance campaign aimed at the newest generation of open-weight frontier models.

While the AI headlines this month have been dominated by model launches and chip geopolitics, inference engines like vLLM are the layer that decides what those models actually cost to run. Version 0.28.0 targets the two families that define the current open-weight frontier — Moonshot’s Kimi-K3 and DeepSeek’s V4 — with a stack of optimizations that push throughput, memory efficiency, and hardware reach in every direction at once.

Decode Context Parallel arrives for Kimi-K3

The headline engineering item is Decode Context Parallel (DCP) support for Kimi-K3 (#50484). Context parallelism — splitting a long sequence’s attention computation across multiple GPUs — has long been standard during prefill, but the decode phase, where tokens are generated one at a time, has resisted the same treatment. Extending parallel decode context to a hybrid-attention MoE model like Kimi-K3 means operators can now spread very-long-context serving across accelerators instead of letting KV-cache memory wall off per-GPU capacity.

DCP is only the spearhead of a broader Kimi-K3 push. The release adds fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), GEMM-RS sequence parallelism (#52079), and combined all-gather operations that deliver 1.5–3x kernel-level speedups (#51070). An adaptive speculative token budget improves DSpark time-to-first-token by roughly 60% (#51725), and optional shared-expert sharding claws back ~17 GiB of memory per GPU (#50912) — headroom that translates directly into larger batches or longer contexts on the same hardware. Kimi-K3 also now runs on AMD ROCm via the V2 model runner (#51653), loosening Nvidia’s grip on serving this model family.

DeepSeek V4 gets end-to-end sparse MLA

DeepSeek’s Multi-head Latent Attention (MLA) — the KV-cache compression trick behind the V-series’ economical long-context serving — gains a major milestone: sparse MLA now works end-to-end, covering plain decode, multi-token prediction (MTP), and DSpark speculative decoding (#51538). Previously operators had to choose between sparse attention’s throughput wins and speculative decoding’s latency wins; 0.28.0 removes the trade-off for V4 deployments.

The DeepSeek work is rounded out by AMD Quark NVFP4 support with an emulation kernel (#47972), reasoning-effort prompt mappings (#50580), sparse top-k metadata kernel optimizations (#52084, #51967), narrowed eager CUDA graph regions (#51430, #52401), and ROCm enablement on gfx11 and gfx950 (#47017, #52212) — the latter bringing V4 serving to AMD’s MI350-class silicon.

Speculative decoding: DFlash2 and confidence scheduling

Speculative decoding — where a cheap draft model proposes tokens that the target model verifies in batches — takes two steps forward. DFlash2 lands with local convolution and a new candidate selector (#52816), while DSpark confidence-scheduled verification (#47808) lets the verifier adapt how aggressively it accepts draft tokens based on measured confidence rather than a fixed schedule. Async scheduling is now auto-enabled for draft models (#48341), removing a knob most operators never knew they should turn.

Tiered KV cache spills to disk

For memory-constrained deployments, KV-cache tiering graduates: 0.28.0 adds disk offloading behind the SimpleCPUOffloadConnector (#49644), out-of-tree secondary tier managers pluggable via module_path (#51007), partial secondary-tier load results (#50321), tiering metrics (#48798), and a canonical CPU layout that makes offload parallelism-agnostic (#48414). The Mooncake transfer engine gains store-group semantics, tenant IDs, and official wheels bundled in the Docker image (#44956, #48069, #51067). Together these changes let a serving cluster treat RAM and NVMe as one latency-hierarchical pool behind GPU memory — the architecture hyperscalers built internally, now available in open source.

Model Runner V2, Rust frontend, and new defaults

The next-generation Model Runner V2 continues its maturation with E/P/D disaggregation (splitting prefill, decode, and encoder roles across dedicated instances) (#38390), weight offloading (#51413), multi-layer MTP KV cache support (#50062), encoder CUDA graphs (#49852), attention-free model support (#52374), and a thinking_token_budget knob (#46727) for capping reasoning tokens — a cost-control feature agent developers have been requesting for months.

The Rust frontend grows up too: a standalone renderer (#50289), multimodal image inference over gRPC (#50368), explicit data-parallel rank routing (#51178), RL lifecycle control (#51316), and protobuf schemas now published to Buf (#51276) for polyglot integration.

Three default changes will be felt immediately on upgrade: max_num_batched_tokens doubles from 8192 to 16384 (#51726), prefix caching switches on by default for Mamba-based models (#50991), and Blackwell CUDA graph capture rises to 1024 (#49390). Operators should re-benchmark rather than assume old tuning tables still apply.

Broader model and hardware coverage

New models include Muse Glimmer (#51655), Ling 3.0 Flash in BF16 with MTP plus FP8 and hybrid MXFP4 variants (#51045, #51265, #52114), Dots3 NOTE native multimodal (#51255), and Interns2mobius (#51149). Qwen3.8 lands on AMD ROCm (#50068), GLM-5.2 gets extended CuTe DSL skinny-GEMM MoE support (#49791), and correctness fixes touch MiniMax-M3 NVFP4 (#48929) and Gemma 4 audio batching (#50958).

Hardware support now spans NVIDIA SM90/SM100/GB10, AMD GFX120x/gfx11/gfx950, Intel XPU (with a wheel added to the release pipeline), CPU (including an MLA backend that lets DeepSeek V2/V3 run on CPU, #49453, plus s390x and Power builds), and ROCm Docker images on torch 2.12 / triton 3.7. A security fix closes a DoS via sample-rate forgery that bypassed the audio decode duration guard (#49948), and docs now explicitly warn that --api-key does not gate every endpoint (#51999).

Breaking changes to watch

Upgraders should note: bitsandbytes support moved to an out-of-tree plugin (#43529), Transformers bumped to 5.15.0 (#51668), the deprecated calculate_kv_scales was removed (#49389), override_attention_dtype was removed (#48684), KV offload tiering metrics were renamed (#52812), and legacy MoE code paths were deleted (#51078). Docker users get an Ubuntu 24.04 runtime base (#51058).

Why it matters

The 2026 open-weight explosion — Kimi-K3, DeepSeek V4, Qwen3.8, GLM-5.2, Tencent’s newly open-sourced 770B-parameter Hy4 — only matters to operators if serving software can extract the economics those models were designed for. vLLM 0.28.0 is a dense, pragmatic answer: parallelize decode, make speculation smarter, spill KV cache down the memory hierarchy, and spread across every accelerator vendor willing to ship a kernel. Seventy-six first-time contributors also signals that the corporate cavalry — from AMD, Intel, and model labs alike — is now contributing upstream rather than maintaining forks.

For teams self-hosting frontier open weights, 0.28.0 is less an optional upgrade than the quarter’s table stakes.