Faster Than It Plays: Video DeltaNet Generates 14 Seconds of 768p Video in 11 Seconds
UC Berkeley, Impossible Inc., and UT Austin researchers bolt a hybrid linear-attention branch onto MiniMax H3, cutting a 14.4-second 768p render from 14 minutes to 11.23 seconds on 8 B200s — with near-lossless quality and a fully open release.
Video generation has a stubborn economics problem: frontier models render slower than the clips they produce play. A team from UC Berkeley, Impossible Inc., and UT Austin just published an answer. On September 6, 2026, the group released VDN-MiniMax-H3 (VDN-H3), a hybrid-attention rework of MiniMax’s open-weight omni-modal video model that generates a 14.4-second 768p clip in 11.23 seconds on eight NVIDIA B200 GPUs — video that arrives faster than you can watch it, at quality the authors describe as near-lossless against the original.
The bottleneck: attention eats the runtime
The starting observation is blunt: on a frontier omni-model like MiniMax H3, softmax attention over long token sequences accounts for more than 85% of total runtime. Softmax attention scales quadratically with sequence length, and video token sequences are enormous — every latent frame contributes a grid of spatial tokens, plus text and audio context.
Large language models already solved an analogous problem by swapping softmax for linear attention, but that trade degrades exactly what video needs: global subject identity, scene layout, and long-range temporal consistency. Pure linear attention forgets; video cannot afford to.
Two branches instead of one
Video DeltaNet’s answer is a hybrid split of the video–video attention into two complementary paths:
- A sliding-window softmax branch handles local detail. Consecutive latent frames are grouped into five-frame chunks (following the original H3 architecture), and each chunk attends to itself, the previous chunk, and the next chunk. This preserves fine-grained local interaction and short-term temporal stability.
- A linear branch handles long-range context. For each frame, a forward state summarizes everything before the local window and a reverse state summarizes everything after — bidirectional, disjoint, no double-counting. Text conditioning is written into both states once at initialization, each scaled by one half, so the summed readout counts the prompt exactly once.
The design adds 4-way boundary anchors — every frame attends to the first and last frames, and vice versa — at a cost of just 3.57% additional attention density, which the team says substantially improves long-range stability. Content-dependent sigmoid gates then calibrate each branch’s contribution into the residual stream.
The core novelty sits in the linear path, which the authors call Video Delta Attention (VDA). Prior frame-wise linear updates (as in SANA-WM) process each spatial token independently against a frozen state; VDA instead solves a joint regularized least-squares problem for the entire frame at once. The closed-form update — S_t = (S̄ + B)(I + A)⁻¹, where A is the weighted key Gram matrix and B the value–key write matrix — turns out to be provably non-expansive for any token alignment, which is why VDA needs no frame-size rescaling, and correlation-aware: repeated evidence in a frame accumulates without cap, while independent directions are not penalized by unrelated tokens.
Plug-and-play: the backbone stays frozen
Crucially for adopters, VDN-H3 does not retrain MiniMax H3. The checkpoint adds the linear attention branch plus two small LoRA adapters — a 50-step variant (4.3 GB) and an 8-step turbo variant (5.1 GB) — that merge into the backbone at inference without touching its weights. The full download is about 82 GB against the 72 GB H3 base.
Training proceeded in three stages: per-layer calibration of each new linear branch (following the Taylor-Calibrate initialization strategy), end-to-end branch adaptation with the softmax path frozen, and finally joint co-adaptation of the linear branch with QKVO LoRA adapters. Notably, the softmax gate stayed frozen until the final stage — a trainable gate earlier would let the optimizer cheat by suppressing the softmax branch rather than learning the new one.
The numbers
The measured stack-up on a single B200, generating the 14.4-second 768p workload:
- Dense MiniMax H3, 50 steps: 13.95 minutes
- VDN-H3 FP8, 50 steps: 5.3 minutes
- VDN-H3 FP8, 8 steps: 51 seconds
Per H3 block, latency drops from 332.5 ms (dense cuDNN) to 192.1 ms with hybrid attention, then 125.3 ms with optimized kernels and FP8 linears — a 2.65× per-layer speedup. The team then shards across eight B200s with Ulysses sequence parallelism (1.62 s/NFE), and improves on it: profiling showed that splitting the two branches onto different GPUs — 5 GPUs on softmax, 3 on VDA — cuts latency a further 13.3% to 1.405 s/NFE. Finally, distribution matching distillation (DMD2), continuing from LarryVRH’s community MiniMax-H3 turbo LoRA, compresses generation to 8 denoising steps.
The end result: 11.23 seconds of compute for 14.4 seconds of video — a 74.5× speedup over the dense single-GPU 50-step baseline, and still 10.7× over dense H3 on eight GPUs. On eight H200s the same clip takes 18.3 seconds; on a single H200, 90.5 seconds. As the model card puts it in one line: the model “generates video faster than it plays.” In practice, that means H3-quality content can be streamed continuously on a single 8-GPU node.
Quality claims are measured, not just asserted: the paper reports generation nearly identical to dense H3, and higher quality with better instruction-following than MiniMax’s own FastH3 fast variant.
Genuinely open — with a licensing landmine
The release is more than weights. The team emphasizes it does not “just open-source the weights”: the optimized inference stack and the training code ship together in the GitHub repository, with kernels built on FlashAttention, Triton, Flash Linear Attention, and FlexAttention, plus support from MIT Han Lab’s Kernel Design Agents for kernel design. The project acknowledges Haoyi Zhu’s “Reflections on Video DeltaNet” as concurrent work.
The catch is the license. VDN-H3 inherits the MiniMax H3 Community License Agreement, whose applicable territory is worldwide excluding the European Union, the United Kingdom, South Korea, and the United States. For the audiences most likely to deploy an open text-to-video model — American startups and European studios — commercial use is not authorized without a separate MiniMax license. Reddit commenters were quick to price the hardware angle too: eight B200s run roughly $40/hour at favorable cloud rates, so “faster than playback” is real-time generation for the price of a small GPU cluster.
Why it matters
VDN-H3 is an architecture study, not a product launch — but it demonstrates two things with unusually clean evidence. First, that hybrid attention is now a practical recipe for video diffusion, not just LLMs: a trained-from-scratch linear branch can carry long-range context at near-lossless quality while softmax shrinks to a local-detail specialist. Second, that the open ecosystem can move faster than the labs that ship the base models: the 8-step distillation builds on a community turbo LoRA, and the entire pipeline — training code included — is reproducible.
The concrete follow-ups to watch: whether the branch-specialized 5/3 GPU split generalizes to other hybrid models, and whether MiniMax — or another lab with a friendlier license — folds a VDA-style linear branch into a future backbone natively. When generation outruns playback, interactive and streaming video applications stop being a latency fantasy. That threshold has now been crossed in public, with receipts.