← All posts / Research

One Brain, Many Bodies: DeepMind's Bet That Gemini Can Jump Between Robots

A rare look inside Google DeepMind's physical AI push shows Gemini Robotics 2 controlling humanoids from feet to fingertips — while robots still fail at dustpans and grape bags.

One Brain, Many Bodies: DeepMind's Bet That Gemini Can Jump Between Robots

In an office full of “galumphing, fidgeting robots,” Kanishka Rao, a principal software engineer at Google DeepMind, is chasing an idea he absorbed as a kid from Star Wars droids and Rosie the Jetsons: helpful robots, sassy optional. The vehicle for that ambition is Gemini Robotics 2, DeepMind’s vision-language-action model for physical AI, and the subject of a detailed new Scientific American report published September 15, 2026 that offers a rare insider’s view of how far the project has actually come — and where it still visibly struggles.

The stakes are enormous. A bullish 2025 Morgan Stanley report projected a humanoid robot market worth upward of $5 trillion by 2050. Elon Musk has made Tesla’s Optimus a primary company focus, BMW has deployed humanoids on a factory floor, and China is packed with robot startups — Unitree is launching a $900 million IPO after rushing several models through U.S. certification about a month before the FCC barred new foreign-built humanoids and robot dogs. An extraordinary amount of money, and now a fenced-off market, rides on machines that still struggle with everyday chores.

From the waist up to feet-to-fingertips

DeepMind unveiled Gemini Robotics 2 in late July 2026. The previous generation controlled person-shaped robots mostly from the waist up. The new model controls whole-body movement — legs, torso, arms, fingers — under a single learned policy, across machines ranging from two-armed research platforms to Apptronik’s Apollo 2 humanoid. In demos, robots fetch snacks on command, change lightbulbs, and tie knots.

The deeper goal is the trick that made large language models powerful: feed a model enough varied data that what it learns in one place transfers somewhere else. Here, that means carrying a skill from one job, or even one kind of robot body, to the next. “We have an existence proof that a general-purpose model can control different robot morphologies because that’s what humans do,” James Marshall, director of the Center for Machine Intelligence at the University of Sheffield, told Scientific American. “If you drive a car, once you’re proficient it feels like an extension of yourself.” Moving from a quadruped to a humanoid or a drone, he cautions, remains “more challenging and data-intensive.”

How the stack actually works

When a Googler sends a toddler-size bot to fetch a bag of popcorn from a nearby lounge, two models cooperate. A higher-level reasoning model, Gemini Robotics ER 2, breaks the verbal request into a series of steps. A lower-level model turns what the robot sees and is told into the movements needed to execute them, updating its predictions of what happens next four to five times every second. ER 2 also watches the task unfold and estimates its progress, classifying each moment into one of five completion ranges — a skill that sounds comically basic until you consider how often a robot can execute perfectly reasonable motions and still fail its overall objective.

The numbers are honest, if sobering. DeepMind reported 57.4 percent accuracy on the progress-classification benchmark — better than the models it tested against, but nowhere near omniscience. In DeepMind’s own tests, an Apollo humanoid completed a countertop-sweeping-into-a-dustpan task just 32 percent of the time. In the observed popcorn run, the researcher eventually got his snack, though he had to pry it “from the robot’s cold, unliving pincer.”

Why manipulation is the hard part

“The key distinction between the backflips and the eggs is this,” Rao explains. “One of them requires you to deeply understand yourself, your own body. The other one requires you to understand the world.” Robots doing backflips or kung fu have mastered gross mobility; folding laundry or making scrambled eggs means mastering everything they touch. Deformable objects — trash bags, brush bristles — fold and bend unpredictably. Grapes are simply fragile.

The deeper deficit is touch. These robots see a lot — they carry more cameras than humans have eyes — but as far as the Gemini model is concerned, they feel literally nothing. A robotic hand can have 22 degrees of freedom and still have almost no sense of what it is touching. When a Gemini-powered robot lifts grapes, no fingertip sensation tells it one is beginning to burst. “The Internet has massive amounts of visual data — gigantic amounts — and it has essentially no tactile data or force data or proprioceptive data,” says Matei Ciocarlie, a Columbia University roboticist who also co-founded robot-hand company Tangent Robotics.

A decade ago, robots were machines repeating motions precise to the centimeter or millimeter, useful in constrained environments like assembly lines. Machine learning changed the equation. Reinforcement learning lets a robot try a task over and over, earning numerical rewards for correct behavior — “a massively powerful paradigm because you no longer have to show it how to do the task,” Ciocarlie says, “but it takes a very long time.” Imitation learning is faster: humans teleoperate the robot or record GoPro-style demonstrations. But it depends on the robot’s digital brain mapping human motions onto its own peculiar collection of joints — the smaller that “embodiment gap,” the better the transfer.

A philosophical reversal, and the Asimov benchmark

One of the most striking notes in the report is how digital AI has inverted an old assumption. Rao, who previously worked on speech recognition and language modeling, expected that a model couldn’t truly know an apple without holding one. “I think I was totally wrong. It’s been the other way around,” he says. “It’s the digital AIs that have really made the physical AI more powerful. It seems like you don’t need to touch all these objects or interact with them or see what they weigh to understand them.” Vanilla Gemini already knows a lot about how the world works, as robotics VP Carolina Parada puts it — what DeepMind is teaching it is “what it feels like to translate it into action.”

Safety work runs in parallel. ER 2 can bring a robot to a safe stop if someone gets too close, and DeepMind released a benchmark called Asimov — named for the author of the Three Laws of Robotics — targeting not just accidental harm but misuse. Gemini provides one safety layer; the bodies supply another, with Apptronik building sensors and safety systems directly into Apollo.

The takeaway

DeepMind’s core bet, in Rao’s words: “If you give [the robot] enough data, intelligence will emerge. That’s the same for language and for motion.” The Scientific American piece ends on a pointed observation — maybe these robots don’t need an inner life, an embodied human understanding of an apple. But as machines that must handle wine glasses, grapes, and dustpans without crushing or fumbling them, “they could use some nerve endings.” Until tactile data becomes as abundant as visual data, physical AI’s frontier will remain exactly where it hurts: at the fingertips.