← All posts / Research

No VLA Required: Stanford's HomeBody Lets GPT Astra Run a Humanoid Directly From a Skill Library

Stanford's Movement Lab skips the learned vision-language-action layer entirely: GPT Astra plus persistent spatial memory drives a Unitree G1 through long-horizon kitchen tasks in an unseen room.

No VLA Required: Stanford's HomeBody Lets GPT Astra Run a Humanoid Directly From a Skill Library

For the past two years, the standard blueprint for humanoid autonomy has been a three-layer stack: a System 2 vision-language model (VLM) for reasoning about what to do, a learned System 1 vision-language-action (VLA) policy for translating perception into commands, and a System 0 whole-body controller for actually moving the hardware. Researchers at Stanford’s Movement Lab (TML), with a Caltech collaborator, have now published a system that asks a blunt question: as frontier VLMs get good enough, do we still need the middle layer at all?

Their answer is HomeBody, a humanoid system that hands GPT Astra — OpenAI’s frontier multimodal model — direct control of a Unitree G1 through a library of five composable motor skills, bypassing the learned VLA policy entirely. In a kitchen the robot has never seen before, the G1 tidies the room across multiple trips and retrieves an object the user only vaguely remembers, with no environment-specific training data and no additional policy learning.

What HomeBody actually does

The demonstration tasks sound mundane, which is precisely the point. In the first, the robot receives an instruction like “clean up all of the coffee bags and put them in the middle, and throw away all of the milk and orange juice cartons that have gone bad.” That requires deciding what to keep versus discard, walking around an island, grasping objects of different sizes, and coordinating repeated trips — a genuinely long-horizon sequence of navigation and manipulation.

The second task is harder in a subtler way. “I forgot my medicine, can you get it for me? Also throw out the bad carton while you are at it.” The medicine is initially out of view. HomeBody must remember where it saw a drawer during exploration, walk over, open the drawer with its right hand, pick the pill bottle out, hand it to the person, and then use its left hand to discard the spoiled carton. Selecting between the two arms lets the robot reach targets on both sides of its body — the kind of whole-body coordination that used to be the exclusive province of learned VLA stacks.

Persistent spatial memory, built by the robot itself

The key enabler is a deployment pipeline with three phases. First, exploration: the humanoid is told “You are a kitchen robot, please explore the space!” and let loose. During this phase it collects 0.5× iPhone video, Intel RealSense D435i camera observations, LiDAR scans processed with SLAM, joint poses, and waypoints chosen by Astra itself. Critically, exploration captures the room from the robot’s own viewpoint — the spatial context is grounded in what the machine can actually see and reach, not in a human-recorded walkthrough.

Second, Real2Sim: Astra acts as a reconstruction agent that builds a digital twin of the room in NVIDIA Isaac Sim from the robot’s own collected data. The FAQ is unusually candid about why this matters — human-recorded video gives rich visual detail, but the robot also needs accurate geometry to navigate, position itself, and reach objects. HomeBody feeds the Real2Sim agent measured SLAM geometry alongside camera observations, producing a twin that is geometrically as well as semantically accurate. At deployment, the G1 localizes in this shared coordinate frame using Super Odometry, with the SLAM map aligned to the reconstruction via iterative closest point (ICP) registration. Ego-camera observations are stored in the same frame with descriptive content, so the robot can return to a recorded position even after objects leave its ego view.

Third, the task itself: Astra receives the instruction and decides where to go and what to act on, with no action-level script.

The architecture: tool calls instead of action tokens

Under the hood, HomeBody replaces the learned action pipeline with a plug-and-play interface that looks a lot like modern agent tool-calling. The VLM selects a skill and a spatial target based on the current ego view, map context, gripper state, recalled observations, and the previous result — then passes the selection through a structured tool call. The skill plans and executes the motion; the VLM never needs to know the skill’s low-level implementation.

The five skills — pick, place, open drawer, pick from drawer, and navigate — share a common interface for targets and execution results, which is what lets the VLM compose them freely at deployment. Each skill hides a respectable stack of classical robotics: a pick call specifies an image point (normalized to 0–1000) and which hand to use; the point prompts SAM 2.1 segmentation with SAMURAI memory selection, Fast-FoundationStereo estimates depth from the stereo images, camera calibration projects the masked geometry into 3D, and the grasp pose is predicted analytically. Arm motion uses a minimum-jerk spline reference with inverse kinematics solved along the path and swept-collision checking.

Error recovery is deliberately layered. If a target shifts in the camera view during approach, visual servoing on the tracker output corrects alignment without bothering the VLM. If a grasp closes on air, the pick skill tries another candidate or adjusts its stance — bounded local retries. Only when recovery is exhausted does the skill return a failure reason to the VLM, which can reposition, pick a new target, or change the plan outright. Walking and manipulation are coordinated through pretrained AMO upper-body-aware lower-body control, with arm and hand commands running at 250 Hz and the AMO policy updating at 50 Hz.

Perhaps the most striking engineering fact: the entire local stack — skills, perception, motion planning — runs on a single Razer Blade laptop with an RTX 4090. Astra runs remotely, sending skill requests down and receiving execution results back.

Why skipping the VLA layer matters

The VLA-optional claim deserves scrutiny rather than hype. Learned VLA policies (of the π0/RT-2/RDT lineage) excel at low-level dexterity and contact-rich manipulation that a library of five parameterized skills cannot cover — HomeBody’s own limitations section concedes the point from the other direction. What HomeBody demonstrates is that for a specific and economically important class of tasks — long-horizon, multi-object, whole-room housework built from mid-granularity primitives — a frontier VLM with persistent spatial memory, accurate localization, and a well-designed skill interface can already do the job without task-specific training.

The trade-offs are real: Real2Sim reconstruction adds setup time and API cost per environment; task length is bounded by hardware endurance (finger-servo overheating during extended operation is called out explicitly); and Astra’s reasoning latency introduces pauses between skills. The skills themselves are also fixed — the vision is that new skills, whether learned policies or classical algorithms, can be added through the shared interface and immediately composed by the VLM.

Still, the timing is what makes this more than an academic demo. The paper lands in a week when agent frameworks, structured tool-calling, and agentic misuse incidents dominate AI news — and HomeBody applies exactly that software paradigm to physical hardware. If frontier VLMs continue improving at their current pace, the “middle layer” of learned policies may end up confined to the dexterity-heavy bottom of the stack, while planning, memory, and error recovery move to models anyone can rent by the token.

The code is available on GitHub, and the kitchen was provided by the Stanford Robotics Center. For a field that has spent two years training ever-larger action models, a laptop, a robot, and a well-prompted frontier VLM is a provocatively lean counter-proposal.