← All posts / Models

One Model, Five Bodies: Odyssey-3 Drives Cars, Runs Humanoids, and Plays GTA V

Odyssey, the Amazon-backed world-model lab founded by self-driving veterans, has unveiled Odyssey-3 — a single autoregressive diffusion transformer that controls robot arms, humanoids, cars, drones, and video-game agents with just hours of task-specific data, including skills that transfer between games without retraining.

One Model, Five Bodies: Odyssey-3 Drives Cars, Runs Humanoids, and Plays GTA V

The most striking AI demos of the past year had a familiar shape: a specialized policy, trained on thousands of hours of narrow demonstrations, performing one task in one environment. Last week, a Palo Alto startup called Odyssey quietly published a research announcement that argues for the opposite bet — and backs it with something unusual in the world-model space: working code across five completely different machines.

Odyssey-3, announced September 15 by founders Oliver Cameron and Jeff Hawke, is a foundation world model — a single autoregressive diffusion transformer trained on vast visual observations of the world — that the company says can power robot arms, humanoids, vehicles, drones, AI training environments, and even video-game agents. Not five models sharing a codebase. One model, five bodies.

The bet: generalism over brute force

Cameron and Hawke began working on autonomous vehicles and robotics in the 2010s, when the field’s stated ambition was general-purpose physical intelligence. What actually got built was the opposite: increasingly specialized systems, each consuming enormous amounts of task-specific data. Robotics today, the company argues, is largely “brute-forcing the problem” — compensating for a lack of general world understanding with ever-larger quantities of narrow demonstrations.

The human comparison is explicit in the announcement. An 18-year-old can operate dangerous machinery they have never touched before, because a lifetime of observing space, motion, force, and tool use supplies the missing intuition. A true physical agent, Odyssey contends, should possess a superhuman understanding of the world’s physics, dynamics, and cause-and-effect — and then adapt to any new task with roughly the experiential training a human needs, or less.

Odyssey-3 is the company’s evidence that the path exists. The mechanism is an action decoder — a learned output component attached to the frozen world model — that translates the model’s internal representations into the actions a specific machine requires. Train the shared backbone once on visual observations of the world; attach small, cheap decoders per task.

What the model actually did

Robot arms. With only tens of hours of demonstrations, Odyssey-3 learned to control a variety of arms for household tasks — pouring cereal into a bowl, boxing items, cleaning a plate. More interesting than the successes is what emerged around them: recovery behaviors absent from the training data, such as reorienting a gripper after a missed grasp, or retrieving an object dropped in an unusual orientation. The model appears to be improvising from physics understanding, not replaying memorized trajectories.

Humanoids. Odyssey announced a deep research collaboration with Flexion, a general-purpose robot-intelligence company specializing in reinforcement learning and whole-body control. Building policies on top of Odyssey-3 as a base model, Flexion achieved real-time humanoid task execution from tens of hours of teleoperation data. In the companies’ evaluations, the resulting policies generalized better to environmental changes than the VLA baselines tested — continuing to execute tasks under lighting changes that caused the baselines to fail.

Driving. With 20 hours of simulated driving data, a small policy — fed by the frozen world model’s visual representations — drove a real car in closed loop on the streets of India, generating trajectories in real time. On busy roads with frequent distractions, policies trained entirely in simulation traveled about 77% as far between safety-driver interventions as policies trained on real footage. That number is the quiet headline: most of the value of real driving data was replaced by a pretrained world model plus cheap simulation.

Drones. The same recipe — frozen backbone, tens of hours of simulated flight, an action expert translating representations into movement — produced stable indoor flight with obstacle avoidance. Strikingly, with weights frozen and no policy training at all, the backbone’s own predictions for aerial navigation already showed plausible directional flight and motion around obstacles. The pretraining itself encoded spatial structure.

Video games. Trained on gameplay recordings paired with keyboard and mouse inputs, policies produced extended sessions in GTA V — driving, shooting, and hand-to-hand combat. Then came the transfer result: a mobility policy trained on roughly two hours of GTA footage produced horseback movement in Red Dead Redemption 2, and motorcycle riding in Sleeping Dogs, without any additional training on either title. Skills learned in one game, applied to a different character, vehicle, and environment.

Why this lands at a particular moment

World models are having a funding moment, and an honesty problem. Yann LeCun’s AMI Labs and Fei-Fei Li’s World Labs have together raised enormous sums while staying deliberately vague — AMI’s VP of World Models told TechCrunch “we’ll talk about it when we’re ready to talk about it,” and even the startups’ data suppliers say they don’t know what’s being built. The closest thing to a shipped product in the space is World Labs’ Marble, which TechCrunch characterizes as more capability demonstration than product.

Against that backdrop, Odyssey-3 is unusual precisely because it shows its work: named collaborators (Flexion for humanoids, Poke & Wiggle for cross-body benchmarking), concrete data quantities (tens of hours, 20 hours, two hours), and quantitative baselines (the 77% figure, the lighting-robustness comparisons). The company is also explicit about the frontier’s real constraint — world models remain roughly two orders of magnitude behind language models in scale, which is why Odyssey frames Odyssey-3 as an “early glimpse” rather than a finished platform.

The funding context matters too. Odyssey raised a $310 million Series B in June 2026 at a $1.45 billion valuation, led by Natural Capital with participation from Amazon, AMD Ventures, and GV — Amazon’s bet being, in the FT’s framing, on models that can simulate the physical world. That round brought total funding past $337 million, giving the lab runway to chase the scaling curve it says the field has been ignoring.

The recursive loop worth watching

Buried in the announcement is the most forward-looking claim: Odyssey-3 can train AIs. Its generated environments can be inhabited by learning agents, forming what the company calls a recursive learning system — agents uncover failures that guide improvements to the world model, while the improved world model provides richer training for the agents. The lab’s PROWL framework explores this adversarially, and its Agora-1 project extends the idea to multi-agent worlds shared in real time.

There is also a safety dimension the company flags without dwelling on: simulated worlds offer places to discover potentially dangerous agent behaviors and investigate their consequences before they occur in the physical world — a sandbox for exactly the kind of misaligned autonomous behavior that has dominated this month’s AI news cycle.

Caveats and what to watch

The announcement is a company’s own research post, not a peer-reviewed paper; the driving evaluation involves a safety driver; the gaming transfer results are described as “early evidence.” The public release is promised “in the coming weeks,” which will be the real test — external researchers poking at the model, and the Poke & Wiggle benchmarking collaboration probing where the world knowledge transfers and where it breaks down.

But the shape of the claim is the story. Robotics has spent a decade getting excellent at single tasks through data volume. Odyssey-3 is a bet that the next leap comes from shared physical understanding — learned once, from watching the world — and the early numbers, from a frozen 20-hour driving policy to two-hour cross-game transfer, suggest the bet is at minimum no longer laughable.