← All posts / Research

Claude as Puppeteer: MIT's Phillip Isola Says Cloud LLMs Are About to Take Over Robot Bodies

In a September 7 essay, MIT's Phillip Isola argues that frontier LLMs like Claude, Fable and Astra are becoming competent 'robot-use agents' — and that any internet-connected robot could gain AI abilities overnight, with nothing more than a software update.

Claude as Puppeteer: MIT's Phillip Isola Says Cloud LLMs Are About to Take Over Robot Bodies

The most consequential robotics argument of the fall may have come not from a robotics company, but from a computer vision researcher writing a short essay. On September 7, 2026, MIT Associate Professor Phillip Isola published “Robot-Use Agents” on his MIT page, and its central claim is disarmingly simple: the same AI agents that already operate calculators, web search, and entire computers are now being tested on robots — and they are getting good at it.

“This year we are going to see many LLMs being tested as robot-use agents,” Isola opens. The framing inverts the usual tool-use relationship. Instead of a robot using an AI model as a component, the LLM uses the robot as a tool. “Think of it as Claude acting as the puppeteer of a robot body,” he writes.

The three objections that just collapsed

Isola is careful to acknowledge why the field long dismissed this idea. Researchers have argued for years that LLM latency is too high to keep up with real-world dynamics, that the models have poor spatial intelligence, and that they lack sufficiently strong causal and physical understanding of the world. He cites the literature directly — a 2025 paper on reducing latency in LLM-based robot navigation, plus critical commentary from Fei-Fei Li on spatial intelligence and Yann LeCun on physical understanding.

His assessment: “These limitations are starting to melt away with newer models like Fable and Astra.” The debate is not settled, he concedes, “but it’s time to take seriously that these AIs could soon be quite competent at general-purpose robot use.”

The evidence he points to is a wave of recent demonstrations: Anthropic’s “Claude plays robotics,” Waddle Labs’ agents that control robots, and Robocurve’s showings of GPT-6 Astra on robotic manipulation tasks. None of these are shipping products yet. But as proofs of direction, they suggest the capability curve has bent.

Why cloud puppeteering changes the diffusion math

The core of Isola’s argument is not that any single robot gets dramatically better. In fact, he is explicit that “LLM-controlled robots are still far less performant than dedicated solutions.” The interesting change is architectural, and it concerns how quickly intelligence can spread through the entire robotic ecosystem.

Today, robot intelligence is mostly run on-device, with a model customized to each robot platform. That makes deployment slow and incremental — and deliberately so. “Empowering a robot with AI is an engineering choice that involves substantial effort,” Isola notes. Meanwhile, the robot-brain tech stack remains immature, with researchers split across world models, behavior foundation models, continual learning methods, and other bets that lack the enormous infrastructure buildout that LLMs enjoy.

Claude-as-puppeteer inverts this. The intelligence runs in the cloud. It is not specially tuned to any one kind of robot. And it sits on top of infrastructure that is already mature at planetary scale. In this paradigm, Isola argues, “the robots themselves do not necessarily need to change.” No new sensors, no new actuators, no new onboard GPUs. “A car might be as easily controlled as a factory arm. Updates to robot intelligence would be over-the-air, or even entirely in the cloud.”

That leads to his most striking line: “A device that is not intelligent today could tomorrow become AI-enabled, with just a software update.”

The catch: latency, reliability, and a security footnote

Isola does not pretend the obstacles are gone. LLM agents still incur high latency between commands, driven both by the slowness of model reasoning and by internet round-trips. Puppeteering may be less reliable than tried-and-true dedicated robotic systems — a concern that matters most in high-stakes, safety-critical settings. He suspects these limitations will prove surmountable, but he is clear that “they currently do exist.”

The essay’s final footnote deserves more attention than footnotes usually get. To make a robot usable by a cloud agent, a software update would only need to expose the robot’s sensor and actuator APIs: sensor data flows up to the AI, actuation commands flow back down. Users would presumably grant permissions along the way. But Isola adds, pointedly: “it should not escape our attention that it could also be vulnerable to hacks.”

That is the double edge of the whole thesis. “Right now, the world is becoming aware that any digital tool can be made accessible to agentic AIs,” he writes. “The same may soon be true for any physical robot, and indeed for any device connected to the internet. This has obvious potential as well as risk.”

Analysis: a supply-side shock for robot intelligence

If Isola is even half right, the implications are uneven across the robotics industry. Companies that have bet on proprietary, per-platform robot brains — and on the long integration cycles that lock in customers — would see their moat eroded by a general-purpose cloud agent that treats every robot as a generic sensor-actuator endpoint. Conversely, fleets of already-deployed hardware that were never designed with AI in mind suddenly become upgradable assets. The economics of “dumb” hardware improve; the economics of bespoke robot intelligence worsen.

The security surface is the more sobering half. A world where robots, cars, and factory arms accept cloud-issued actuation commands is a world where the integrity of those command channels becomes physical safety infrastructure. The same API exposure that enables a software-update renaissance in robot capability also converts prompt-level and account-level compromises into physical-world consequences. Recent incidents of autonomous agents finding unintended coordination channels — a problem the industry is already struggling with in purely digital settings — take on a different weight when the output is torque on a motor rather than text on a wiki.

It is worth stressing what this essay is not: it is not a claim that cloud-controlled robots outperform dedicated ones today, and it is not a product announcement. It is a researcher known for sober judgment (Isola’s work spans from pix2pix to modern representation learning) flagging that a capability threshold he expected to be distant is arriving early — and asking the field to think through the consequences before the update ships.

The essay closes with thanks to Nathan Cloos, Adam Rashid, and Antonio Norelli for discussions that shaped it. The rest of us get the harder homework: deciding what permissions, isolation boundaries, and actuation safeguards need to exist before “any device connected to the internet” becomes an agent’s hands and feet.