Skild AI's S1 Learns 10-Minute Robot Tasks From a Single Video — No Fine-Tuning
Skild AI's S1 robotics foundation model executes never-before-seen manipulation tasks up to 10 minutes long from one video demonstration — 66% success on unseen tasks with zero fine-tuning, 7x better than language prompting.
Robotics has been stuck in what Skild AI calls “the BERT era.” For years, the field could learn complex behaviors with deep neural networks, but every new task demanded fresh data collection — hours of teleoperation — and a fine-tuning run to produce a narrow specialist policy. On August 25, 2026, the Pittsburgh-based startup published the first results from S1, its flagship robotic foundation model, and made a claim that directly attacks that paradigm: show S1 a single video of a task, short or long, seen or unseen, and the robot executes it. No fine-tuning, no post-training, one set of weights behind every example in the announcement.
What S1 actually does
S1 is built from the ground up as an in-context learner. The task specification is not language — it is a video demonstration that enters the model’s context window. The demonstration may come from a different scene, viewpoint, or even a human demonstrator rather than a robot, and the policy must implicitly infer the demonstrator’s intent, work out the functional correspondences between what it sees and its own grippers, and track task progress to predict appropriate actions.
The headline capability is long-horizon, out-of-distribution learning. Skild demonstrated four tasks S1 was never trained on: plant potting, pancake cooking, pour-over coffee making, and kit assembly. Each runs for up to ten minutes, spans dozens of manipulation steps, and is driven entirely by one visual demonstration. In the pancake task, the flipping motion itself was a completely novel behavior learned at test time — digging into soil to make room for a plant and pressing a coffee filter into a funnel were similarly new primitives composed on the fly.
The deployment timeline the company published is a quiet rebuke of conventional robot-learning workflows. For the plant-potting task: props arrived at the office at 8:54 PM, the scene was set up by 9:16 PM, a single egocentric human video was recorded at 9:22 PM, and S1 was executing the task autonomously on hardware by 9:27 PM. Eleven minutes from demonstration to autonomous execution, with most of the elapsed time spent moving furniture.
The numbers behind the claim
Skild ran a controlled study comparing in-context (ICL) demonstration prompting against the language prompting popularized by conventional vision-language-action (VLA) models. Both variants used identical data, architectures aside from the prompt embedding, and compute, trained on datasets scaling from 1,000 to 100,000 hours.
On tasks drawn from the pre-training distribution, the two approaches trade blows — at 1k hours the language-conditioned policy actually led, 53% to 43%. But on unseen tasks the gap explodes. Language prompting crawled to a 9% success rate at 100k hours; the ICL model reached 66% — roughly a 7x improvement. Skild attributes this to two structural advantages of demonstrations over instructions: a novel action primitive like flipping a pancake has no grounding in language tokens, and language is often too coarse to specify an unprecedented chaining of known skills, while a demonstration spells the composition out directly.
Expressed differently: a single demonstration in context is worth roughly 380 post-training teleoperation episodes. A conventional policy eventually passes S1’s zero-shot performance — reaching 86% success after 2,000 demonstrations — but it has to be trained on every task, every time.
Robustness results point the same direction. Under the hardest perturbation tier Skild tested — where half the robot’s actions must execute with the opposite arm — the language-prompted VLA degraded up to three times as much as the ICL policy. And emergent behaviors show up unprompted: S1 completes tasks after objects are slid away, swapped, or relit mid-execution; it recovers from its own failures on out-of-distribution tasks like assembling a skateboard wheel; and when the demonstration waters a plant with a watering can but only a cup is available, S1 uses the cup. Most strikingly, S1 sometimes corrects its teacher — when a demonstrator drops an egg and makes a mess, S1 performs the same step with a controlled motion, treating the demonstration as a specification of the goal rather than a trajectory to replay.
Why this matters now
The direct comparison Skild invites is with Generalist AI’s GEN-1.5, which this blog covered on August 22. Skild’s own framing is pointed: concurrent approaches to manipulation ICL “largely cover tasks which are short-horizon or already present in the pre-training distribution.” GEN-1.5’s one-shot demos ran 3–12 seconds. S1’s unseen tasks run up to ten minutes — the first time, Skild claims, a robotics foundation model has shown in-context learning on extremely long-horizon tasks never seen during pre-training.
The strategic logic follows the language-modeling arc from BERT to GPT-3. Understanding-only models still needed fine-tuning for every application; the pivotal transition to ChatGPT was driven by in-context learning, where a novel concept enters through the prompt and weights stay frozen. Skild argues robotics is at the same inflection: pre-training’s main purpose — “if not the only one” — is to enable that shift, because when post-training data is dense enough, even from-scratch policies match post-trained foundation models, eroding the value of pre-training itself.
Underneath the results is a data engine scaling all sources at once, because no single source wins on hardware proximity, diversity, and scalability simultaneously — teleoperation is closest to the robot but scales worst; egocentric video scales best but has the largest domain gap. Skild says it spends three dollars on quality control for every dollar of data collection, screening every datapoint for precision, coherence, and annotation fidelity.
The company’s commercial position makes the claims harder to dismiss. Skild — backed by Amazon, SoftBank, and reportedly Nvidia, with a $1.4B round valuing it above $14B — already runs models on Foxconn assembly lines building Nvidia’s Blackwell servers in Houston, and states S1 “is already at work with our commercial partners.” Fast deployment feeds the flywheel: tasks set up in minutes generate real-world robot interactions that flow back into pre-training. The company followed LocoFormer (locomotion ICL, September 2025) through first in-domain ICL results (February 2026) to S1 flipping its first pancake in May.
This is the first post in a promised series; future installments will detail how S1 is trained. The caveats deserve weight — success rates are self-reported on internal benchmarks, human intervention recovers failures so all steps get graded, and the headline comparisons come from the company’s own lab. But the direction is clear: the robot-learning field is converging on the same conclusion the LLM field reached in 2020, and the cost of teaching a robot a new task is collapsing from days of teleoperation toward minutes of video.
Sources
- [1] https://www.skild.ai/blogs/s1
- [2] https://techmeme.com/260825/p41
- [3] https://www.linkedin.com/posts/pathak22_announcing-s1-our-new-foundation-model-that-activity-7498066901005172736-XNlM
- [4] https://digg.com/tech/rbnvscse
- [5] https://app.dealroom.co/news/note/skild-ai-introduces-s1-a-foundation-model-that-learns-tasks-from-one-example