← All posts / Research

A World You Can Type Into: SeedLeap's Zing-0.5 Runs a Playable 5B World Model at 24 FPS for $0.009 a Minute

Chinese startup SeedLeap open-sources Zing-0.5, a 5B autoregressive world model that streams a keyboard-navigable, text-editable generated world at 832×480 and 24 FPS — scoring 81.0 on WBench Navigation with weights, code, and serving stack all public.

A World You Can Type Into: SeedLeap's Zing-0.5 Runs a Playable 5B World Model at 24 FPS for $0.009 a Minute

The most interesting question in generative media right now is not “how good can a video look?” but “what happens when the video listens to you?” This week, a 16-person Chinese startup called SeedLeap.ai (涌跃智能) published a technical report that pushes hard on the second question. Zing-0.5, released September 15 on arXiv (2609.17909), is a 5-billion-parameter autoregressive world model built for what the team calls playability: you steer with the keyboard, you intervene with text, and the world keeps rolling — one continuous stream, no restarts.

And in a field where frontier results are increasingly locked behind APIs, SeedLeap shipped everything: model weights on Hugging Face, inference code and a Zing-SGLang serving implementation on GitHub, and a technical report detailed enough to reproduce the training pipeline. For anyone tracking the open interactive-video race, this is the reference point of the week.

What “playable” actually means here

Most interactive world models so far have picked one control axis. Navigation-focused systems — the line of work descending from Genie-style keyboard control — let you move through a generated scene, but the text prompt is fixed at the start. Language-conditioned systems let you change what happens, but as a separate mode from movement. Zing-0.5’s contribution is combining them in a single continuous session: WASD keys (plus IJKL for view changes) drive movement and camera, while online text instructions fire mid-stream to change subjects, scenes, or events — and generation continues from the existing visual context.

The demo cases from the paper show the range. A user pilots a boat across a sea of clouds toward distant palaces; mid-voyage, a text instruction ignites a silver-flowered tree, and golden sparks shower down inside the same forest clearing. In another session, first-person hands reach toward a candy house that melts under them. In a third, the viewpoint follows a character through a bedroom mirror into a garden, and the exploration continues outdoors. Subject-directed changes work too: one instruction invokes lunar light and the character rises into flight; another makes a flying mount breathe fire — all while keyboard navigation keeps running.

The three technical ideas that make it work

1. Unified action and text conditioning, with magnitudes. Zing-0.5 builds on Alibaba’s Wan2.2-TI2V-5B video backbone, retaining its video autoencoder and diffusion Transformer, then adds a lightweight action encoder — just 3.68M parameters, about 0.074% of the 5B backbone — that injects frame-aligned control features into visual tokens before the DiT stack.

A subtle but consequential choice: instead of converting keyboard inputs into camera-pose increments (the common scheme inherited from camera-trajectory control), Zing conditions directly on the keys, each augmented with a continuous magnitude. The team argues that pose-based schemes are fragile — if generated motion drifts from the conditioning trajectory, the discrepancy compounds over long rollouts and can break visual coherence, a problem made worse when prompt updates themselves shift the viewpoint. Direct conditioning on native actions keeps the representation identical between training and inference, and the continuous magnitudes preserve fine-grained intensity that discrete labels discard. Notably, they also downsampled the ocean of W-key-only forward clips in the training data so turning and combined controls weren’t drowned out.

On the text side, prompts are aligned to temporal intervals rather than applied globally: each segment of frames attends only to the caption describing that segment, which is what allows mid-stream instruction changes to take effect locally.

2. Event-scale supervision via teacher–student distillation. The training pipeline is a four-stage progression from the pretrained bidirectional video model to a few-step causal generator. The clever part is the split between two branches trained from the same adapted checkpoint: a segment-level teacher that jointly denoises an entire prompt interval with bidirectional attention (it can see the whole event, and since it only works during training it has no latency budget), and a block-level generator that must produce video incrementally in four-latent-frame blocks to respond to inputs with low latency.

The final stage closes the train/inference gap with distribution-matching distillation (DMD): the student trains on its own rollouts, scored by a frozen real scorer and a learned fake scorer both initialized from the teacher branch. The team adapted the two-pass procedure of Self Gradient Forcing — no-gradient rollout first, then packed gradient computation — to keep memory tractable on long, variable-length samples. They also mixed Data-Forcing Distillation (DFD) throughout the DMD stage rather than as a post-training afterthought, which they report reduces fluctuations in motion magnitude and prevents severe visual degradation. The end result: four denoising steps per block, with guidance baked into training so inference needs no second unconditional forward pass.

3. Economics as a first-class constraint. This is where the report gets unusually concrete. On a server with 8 RTX 5090 GPUs, Zing-0.5 serves 8 independent 832×480 streams at a client-visible 24 FPS (unpaced steady state: 24.63 FPS, leaving real-time margin). Each GPU hosts one complete replica — DiT, bounded causal KV cache, and a local TAEHV decoder — deliberately avoiding cross-GPU latent transfer. Four denoising steps produce 4 latent frames, which decode to 16 video frames; the native VAE is only needed for the initial image, with the much lighter recurrent TAEHV decoder handling the rest, and H.264/fMP4 fragments flowing through a bounded ring buffer that favors the live edge under backpressure.

The bottom line the team publishes: an estimated server rental cost of roughly $0.009 per stream-minute. If playable generated worlds are ever going to be a product category rather than a demo genre, that number — not the benchmark score — may be the one that matters most.

How it scores

On the 158-case Navigation split of WBench, generated at 1248×704, Zing-0.5 achieves an overall score of 81.0 with 88.5 consistency — tied on average with JoyAI-Echo-1.5’s 4-step variant and ahead of HiDream-O1-World (80.9), Alaya-EVOKE 3-step (80.8), and LingBot-World v2 fast (79.4), while sitting just behind JoyAI-Echo-1.5’s bidirectional configuration (81.6). Physical plausibility is Zing’s strongest sub-score at 73.8, the best in the selected leaderboard table. All this from a 5B model — far from frontier scale.

The honest limitations section is the best part

What elevates the report above the usual launch-document genre is Section 6, where the team lays out exactly why this is not yet a game engine. Their argument is worth quoting the shape of:

  • Meaningful interaction requires lasting consequences. The demos show visible responses to instructions, but the paper is candid that nothing establishes whether consequences persist. “An object moved aside should remain displaced when the user returns; a rule learned through one interaction should still apply to the next.”
  • Visual history does not fully specify world state. Zing generates from retained visual context, text, and actions — it tracks no explicit entity states or transition rules. An object can leave the frame without ceasing to exist; an event can end while its consequences should still bind. The bounded KV cache limits access to earlier evidence, and autoregressive generation means one bad continuation becomes context for the next.
  • Persistent worlds need architectural support. Their conclusion is that persistence itself must become a design objective — future systems should be evaluated on whether users can change something, leave it behind, and later continue from the consequences of that change.

This is the correct diagnosis of the gap between “impressive interactive video” and “a world,” and the fact that it comes from the team that built the model, in the launch paper itself, is exactly the norm open research should normalize.

Why it matters

The interactive world-model space has been heating up all year — JoyAI’s Echo, Alibaba’s H3-World and HappyOyster, HiDream’s O1-World, Alaya’s EVOKE, LingBot-World — and the WBench Navigation leaderboard is becoming the arena where the open contenders compare. Zing-0.5’s combination of joint control, real-time economics, and a genuinely open release (weights + code + serving stack) sets a high bar for what the next entrant has to clear.

There is also a strategic reading. The Wan2.2 backbone is Alibaba’s open model; SeedLeap’s contribution demonstrates how much mileage a small team can extract from the open ecosystem — a 16-person startup publishing a competitive entry in a race where the other names are major labs. The open-weight flywheel, so far mostly a story about language models, is now visibly turning for interactive video.

For developers, the $0.009-per-stream-minute figure is an invitation: a playable generated world costs about half a cent for a minute of play. Experimentation that was economically absurd a year ago is now a weekend project on rented 5090s. Expect the strangest applications of this model to come from people who never would have gotten permission to build on a closed one.