No Physics Engine Required: D-Robotics' Uranus Generates Robot Simulation Frame by Frame
A diffusion model that skips torque integration entirely — Uranus turns joint trajectories into multi-view video at 24 FPS, spanning five robot embodiments and 2,400+ hours of training data.
Every robot learning team eventually collides with the same wall: physics simulators are cheap to run but poor at looking real, and the real world is photorealistic but catastrophically expensive to run experiments in. This week, D-Robotics AI Lab published a research report that stakes out an aggressive middle position: throw out the physics engine for visual simulation altogether, and let a diffusion model generate what the cameras would see.
The system is called Uranus, described in the paper “Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI” (arXiv:2609.24815, revision 3 posted September 23, 2026, from a team led by Wenkang Qin and Yukun Zhou, with Noah Shen, Jisong Cai, Dongxiao Mao, Baicheng Li, Yue Zhang, and Wei Sui). It surfaced on Hugging Face’s daily papers list on September 24 and has been circulating through the embodied-AI community since.
What Uranus actually is
Uranus is a joint-trajectory-conditioned autoregressive diffusion model. The inputs tell a complete story on their own: an initial set of multi-view observations, calibrated camera intrinsics and extrinsics, a robot embodiment description in standard MJCF or URDF format, and four future joint configurations expressed as qpos. The output is the next synchronized multi-view observation — what the robot’s cameras would actually see if it moved that way.
Each autoregressive step generates one latent frame, which decodes into four consecutive RGB frames per camera, with the four joint-position samples aligned one-to-one to those four timestamps. Because the generated observation is fed back as context for the next prediction, there is no fixed rollout horizon: the simulation can run open-ended, one small chunk of motion at a time. After inference optimization — including disaggregated DiT and VAE decoding plus KV-cache-centric state persistence — the team reports 24 FPS generation.
The contrast with a conventional simulator is architectural, not incremental. A physics engine integrates torques, velocities, and contact forces forward in time, then renders the result. Uranus never integrates anything. It learns the visual consequences of kinematic trajectories from data, which means contact-rich scenes, deformations, and cluttered backgrounds come for free if they appear in the training corpus — and behave wrongly if they don’t. That trade is the whole bet.
Trained on 2,400+ hours of real robot data
The training corpus is assembled from AgiBot World Beta (141,355 episodes, 257.4 million frames, roughly 2,383 hours at 30 FPS), AgiBot World 2026, DROID, and RoboChallenge Table30 v1 and v2 — spanning single-arm, dual-arm, and humanoid platforms. To handle this variety without architecture changes, Uranus decouples embodiment-specific kinematics from the learned representation: each robot body, whether UR5, ALOHA, AGIBOT, ARX5, or DROID, is converted through online forward kinematics and camera projection into a shared image-space skeleton representation. The end-effector marker encodes task-relevant detail directly — its radius represents gripper opening, and its color is computed from spherical harmonics of the relative pose between the end-effector and the camera.
The numbers that matter
Three evaluation threads stand out:
Long-horizon stability. On the WorldOlympiad long-horizon interaction benchmark, the authors measure degradation over extended rollouts. During the final 140–160 second interval, Uranus holds approximately 12.3 PSNR and 0.45 SSIM on the head view, against 7.1 PSNR and 0.20 SSIM for GE-Sim 2.0, the closest contemporary comparison. Every generative simulator degrades as errors compound; the claim here is that Uranus degrades more slowly.
Task-OOD protocol. In a nod to growing concern that embodied-AI benchmarks quietly train on their test sets, the authors split 30 tasks into 24 for training and 6 for held-out evaluation, with a designated test case, and report both in-distribution and out-of-distribution results. Qualitative OOD analyses cover generalization across unseen scenes, trajectories, tasks, camera motions, and robot embodiments.
Closed-loop policy consistency. Perhaps the most consequential experiment evaluates whether Uranus can serve as a policy evaluator: several policies were run in closed loop inside Uranus and in the RoboTwin clean environment, with per-policy mean success rates compared across the two. If simulated success rates rank real-world success rates, a learned simulator becomes a cheap proxy for hardware trials — the exact use case roboticists would pay for.
Read the fine print
The paper is unusually candid about where the approach falls short. The authors explicitly identify remaining limitations in precise contact dynamics and object-state transitions under challenging distribution shifts — precisely the properties a physics engine guarantees by construction. The report publishes no headline win-rate of policies trained in Uranus against policies trained in a physics simulator, which is the transfer experiment the field is still waiting for. And the arXiv history is itself a small drama: version 1 landed September 21, version 2 was withdrawn a day later, and version 3 restored the full 20 MB paper on September 23.
The release includes code and model weights, which is what separates an infrastructure claim from a demo.
Why this matters now
Uranus arrives amid a crowded week for learned world models — Black Forest Labs shipped the open-weights FLUX 3 Action the same week, and the paper itself notes concurrent releases touching the same stack. The shared direction is unmistakable: if video generation models can predict what happens next at interactive frame rates, the bottleneck for robot learning shifts from collecting real data to trusting generated data. Learned simulators like Uranus are the infrastructure bet that the trusting will come.
The open question is the one the paper itself poses in its closing sections: whether policies trained on generated frames transfer to real hardware at rates that justify replacing at least part of the physics-engine pipeline. Until a third party runs that head-to-head comparison, 24 FPS proves Uranus streams — not that it simulates. But as a statement of where simulation infrastructure is heading, it is one of the clearest artifacts yet published.