Skild AI's S1 Learns 10-Minute Robot Tasks From a Single Video
Skild AI's S1 robotics foundation model executes long-horizon tasks it never saw in training from one video prompt — 66% success on unseen tasks versus 9% for language-prompted VLAs.
Pittsburgh-based Skild AI has released S1, a robotics foundation model built from the ground up as an in-context learner: show it a single video demonstration of a task — short or long, seen or unseen during training — and the robot executes it, with no fine-tuning and no post-training. In demos, S1 performed tasks running up to ten minutes long that never appeared in its pre-training data, including pancake cooking, pour-over coffee brewing, plant repotting, and kit assembly — each driven by one egocentric human video as the entire task specification.
Why this matters
Robotics has been, in Skild’s own framing, “stuck in the BERT era.” Modern robot learning can acquire complex behaviors, but deploying a policy for a new task still demands hours of teleoperated data collection plus a fine-tuning run — a cycle that repeats for every task. Worse, research cited by Skild (Oh et al., 2026) shows that when post-training data is large enough, policies trained from scratch can match a fine-tuned foundation model, calling into question what pre-training buys you at all.
Skild’s answer mirrors the leap from BERT to GPT-3 in language modeling: the entire point of pre-training is to enable in-context learning (ICL) — acquiring a new behavior from one or a few examples, at inference time, without touching the weights. S1 is the first robotics foundation model to demonstrate ICL on extremely long-horizon tasks (up to 10 minutes) that were completely absent from pre-training.
How S1 works
The training recipe is conceptually simple: pre-train on episodic data where the task is specified only through an in-context video demonstration. Because that demonstration may come from a different scene, viewpoint, or embodiment, the policy is forced to implicitly infer the demonstrator’s intent, functional correspondences, and task progress in order to predict correct actions. In meta-learning terms, pre-training is the outer loop that teaches the policy how to learn from context; at inference time, the demonstration drives the inner loop without changing any weights. One set of weights produced every example in Skild’s release — no per-task specialists.
Task diversity and scale are what drive the capability. In the high-diversity regime, scene ambiguity must be resolved by attending to the context, which incentivizes the model to learn how to learn from demonstrations. S1 is trained on NVIDIA AI infrastructure, and Skild deliberately scales all data sources rather than betting on one:
- Teleoperation — high hardware proximity, low diversity, scales worst
- UMI-style handheld capture — moderate on all three axes
- Egocentric video — highest diversity and scalability, largest domain gap
- Simulation — scalable, but limited diversity
Skild also notes an unusual cost structure: for every dollar spent collecting data, three are spent on quality control, with every data point screened for precision, task coherence, and annotation fidelity.
The numbers
In a controlled study comparing ICL-style demonstration prompting against the language prompting used by conventional VLAs, Skild trained both policy types on identical data, architectures, and compute, from 1k to 100k hours:
- Seen tasks: at 1k hours, language prompting actually led (53% vs 43% for ICL). But ICL overtakes VLAs as data scales, which Skild attributes to language ambiguity — an instruction admits many valid executions, while a demonstration specifies a mode.
- Unseen tasks: the real divide. Language prompting creeps up to 9% success at 100k hours; S1’s ICL reaches 66% — a 7× improvement at identical training scale.
- Demonstration efficiency: a single in-context demonstration is worth roughly 380 post-training episodes for a VLA. Post-training does eventually win — 2,000 demonstrations got a fine-tuned VLA to 86% — but collecting 380 demonstrations for long-horizon tasks means 50–100 hours of teleoperation. Teaching S1 something new takes minutes.
The deployment-speed claim is vivid: Skild’s plant-potting experiment started at 8:54 PM, when soil, a pot, a watering can, and a plant arrived at the office. After setup, one human demonstration was recorded at 9:22 PM, and S1 was executing the task autonomously at 9:27 PM — 11 minutes from demonstration to autonomous execution.
Emergent behaviors
Beyond raw success rates, Skild documents behaviors that emerged without targeted data collection:
- Perturbation robustness — objects slid away mid-execution, swapped, or lighting changed (none of it shown in the prompt); S1 completed tasks anyway.
- Mistake recovery — when it fails, it retries rather than blindly proceeding, even on out-of-distribution tasks like assembling a skateboard wheel.
- Common sense — prompted with a watering can but given only a cup, S1 used the cup; asked to fill an already-nearly-full glass, it just topped it off.
- Demonstration correction — when the human demonstrator dropped an egg and made a mess, S1 performed the same step with a controlled motion. It treats the demonstration as a specification of the goal, not a trajectory to reproduce.
Under a five-level distribution-shift ladder (from 15 cm object shifts up to forcing half the actions onto the opposite arm), the language-prompted VLA degraded up to 3× more than the ICL policy — VLAs handle small pose variations but collapse when novel motions are required, while an in-context policy can simply be prompted with an execution plan that fits the deployment scene.
Context and what’s next
Skild AI, spun out of Carnegie Mellon research, raised a $1.4 billion Series C led by SoftBank in January 2026 at a valuation above $14 billion — reported as the largest robotics-AI funding round on record — after a $135M Series B at a $4.5B valuation just seven months earlier. The company’s “omni-bodied” thesis is a single brain across robot morphologies, and its roadmap milestones tell the story: LocoFormer (in-context locomotion) in September 2025, first in-domain ICL manipulation results in February 2026, the first out-of-distribution pancake flip in May 2026, and the S1 release in August 2026.
S1 is already deployed with commercial partners, and Skild says rapid task setup feeds a data flywheel: minutes-long deployments mean real-world robot interactions can be bootstrapped straight back into pre-training. This post is the first in a series; follow-ups will detail how S1 is trained.
The competitive landscape is moving in the same direction — concurrent work like RoboTTT (Jiang et al., 2026) and Generalist AI’s one-shot embodied models — but those efforts remain constrained to short horizons or in-distribution tasks. If Skild’s scaling curves hold, the economics of robot deployment flip: the bottleneck stops being data collection for each task and becomes the quality and diversity of pre-training data overall. That is precisely the transition that turned language models from research prototypes into general-purpose tools — and the reason a single video, rather than a thousand teleop sessions, may soon be the way you teach a robot.