← All posts / Models

World Labs Unveils Atlas: An Omni World Model That Generates, Reconstructs, and Simulates Any World

Fei-Fei Li's World Labs introduced Atlas on September 1 — a multimodal autoregressive diffusion transformer pretrained from scratch to natively handle text, images, video, and 3D, generating minute-long 1440p videos and beating specialist models at 3D reconstruction.

World Labs Unveils Atlas: An Omni World Model That Generates, Reconstructs, and Simulates Any World

On September 1, 2026, World Labs — the spatial intelligence company founded by AI pioneer Fei-Fei Li — pulled back the curtain on Atlas, its next-generation “omni world model.” Where a large language model predicts the next token and a video model predicts the next frame, Atlas is built to do something more ambitious: generate, reconstruct, and simulate entire worlds, natively operating across text, images, video, and 3D within a single unified architecture.

The announcement is the company’s biggest research milestone since its first image-to-3D-world demo in December 2024, and it arrives at a moment when “world models” have become one of the most hotly contested frontiers in AI. World Labs raised a $1 billion round in February 2026 — including $200 million from Autodesk — to bring world models into professional 3D workflows, acquired robotics-simulation startup SceniX, and shipped its first commercial product, Marble, in late 2025. Atlas, the company says, “will power future versions of Marble and other products.”

What Atlas Actually Is

Under the hood, Atlas is described as a multimodal autoregressive diffusion transformer, pretrained from scratch. Each of those three words carries weight:

  • Multimodal — Atlas natively processes text, images, camera poses, and 3D depth maps, with videos represented as sequences of images. Crucially, every image and depth map is conditioned on an explicit camera pose, making spatial control a first-class architectural citizen rather than an afterthought.
  • Autoregressive — Like an LLM, it operates on sequences, generating each output element one at a time conditioned on everything that came before. Each task is simply a different kind of sequence.
  • Diffusion — Atlas is a rectified flow model that denoises its way to outputs, excelling at high-dimensional continuous data and able to trade speed against quality by varying denoising steps.

The unifying idea is what World Labs calls the spatial context: all inputs are combined into a shared representation in which each image is grounded at a 3D position in space. Atlas generates what comes next, staying geometrically consistent with everything it has seen — and imagining what lies beyond it. The company notes the design borrows deliberately from both worlds: KV-caching, cache-aware routing, and disaggregated serving from the LLM side; diffusion distillation, classifier-free guidance, shifted noise schedules, and modern VAE design from the video-model side.

Four Capability Pillars

Camera-controlled generation. Atlas generates images and videos from one or more reference images with what the company calls pixel-perfect camera control — precise camera geometry as a native input type, not coarse text instructions like “pan left.” Output reaches up to one minute of video at 1440p. In a striking demo, the model builds a complete explorable scene from a single photo, correctly inferring the back side of a robot and guessing a grassy lawn sits next to a pool.

Spatial reconstruction. Atlas reconstructs real-world scenes from as few as one to three input images — no capture rigs, no dense views — producing both novel-view frames and explicit 3D outputs (point clouds and 3D Gaussian splats). Feed it more images and it imagines less: with over a hundred inputs it faithfully recreates full real-world environments. In one example, it rebuilt Stanford’s Main Quad from two to twenty-five ground-level photos, then generated aerial flyover paths far above campus.

Space-time simulation. Atlas doubles as a world simulator. The VFX showcase: feed it footage from as few as three ordinary phones on tripods, and it can freeze time and reframe the shot from impossible angles — a “bullet time” studio in a backpack. The robotics showcase is bigger: from a casual 24-frame cell-phone video, Atlas reconstructs a space in 3D and then generates the RGB and depth data a simulated robot’s body-mounted cameras would observe along any path — with variations in objects, positions, lighting, and background for large-scale Real-to-Sim training data covering rigid, articulated, and deformable objects.

Image generation. Although not its focus, Atlas is a capable image generator, following complex prompts, rendering text, producing 360° panoramas, and spanning a wide variety of visual styles.

The Numbers

On camera-controlled generation, third-party human raters judged Atlas head-to-head against leading video models given the same input image and camera path. The results were lopsided: raters preferred Atlas over MiniMax H3 75% of the time, Gemini Omni Flash 81%, Happy Horse 1.1 86%, FLUX 3 93%, and Seedance 2.5 94% — with the advantage growing as camera trajectories become more complex.

On 3D reconstruction from sparse posed views, Atlas — a generalist — beat the best specialized open-source reconstruction models with a mean absolute-relative pointmap error of 25.3 (×10⁻³) averaged across seven standard benchmarks, ahead of Pi3X at 28.7, VGGT-Ω 1B at 34.7, Depth Anything 3 at 36.4, π³ at 39.3, and MapAnything at 47.7. Per-dataset wins included DTU (8.6 vs 11.1), ETH3D (9.3 vs 18.7), ScanNet (12.4 vs 15.7), and 7-Scenes (37.8 vs 39.3).

Scaling Is the Point

World Labs also published scaling evidence: it trained a series of Atlas models of increasing size and training compute, and found that “each new level of compute unlocked new model capabilities.” The company says it expects the trend to hold as it continues scaling — a pointed signal that Atlas is a foundation to be built upon, not a one-off demo.

Why It Matters

The world-model race has quietly become the third front of the AI boom, after chatbots and coding agents. Robotics companies need simulators that don’t require hand-built digital twins; game and film studios want controllable world generation; and the enterprise 3D world — CAD, BIM, digital twins — is exactly the market Autodesk paid $200 million to reach through World Labs.

Atlas’s bet is that neither the pure-LLM path nor the pure-video-model path gets there alone. By making camera poses and depth native input types and grounding everything in a shared spatial context, World Labs is arguing that spatial intelligence — Fei-Fei Li’s long-running thesis — requires architectures designed for space from the first layer, not bolted on afterward. The benchmark numbers suggest the generalist already outperforms the specialists at their own games, which is precisely the pattern that preceded LLMs swallowing NLP.

Atlas is available now in early access with select partners, and interested teams can request access through World Labs. If the scaling curves hold, today’s Atlas may be remembered as the moment world models stopped being demos and started becoming infrastructure.