← All posts / Research

Skild's S1 Learns Robot Tasks From a Single Video — No Fine-Tuning Required

Skild AI's S1 executes unseen 10-minute manipulation tasks from one human video demo, scoring 66% success versus 9% for language-prompted VLAs — the 'BERT-to-GPT-3 moment' for robotics.

Skild's S1 Learns Robot Tasks From a Single Video — No Fine-Tuning Required

Robotics has spent the past decade stuck in what Skild AI calls “the BERT era.” Deep networks can learn complex behaviors, but every new task still demands hours of teleoperation data collection and a fine-tuning run before a robot can do it. On August 25, 2026, Skild AI released S1, a robotic foundation model built to break that pattern the same way ChatGPT broke it for language: through in-context learning. Show S1 a single video of a human performing a task — a task it has never seen during training — and it executes, for up to ten minutes of continuous manipulation, without a single gradient update.

The BERT-to-GPT-3 Moment for Robots

The analogy Skild draws is deliberate and precise. BERT-class language models were strong understanders, yet every downstream application still required additional data collection and fine-tuning. The leap to GPT-3 and ChatGPT came from in-context learning (ICL): users introduce novel concepts entirely through the prompt, and the model responds without touching its weights. Robotics, Skild argues, has been stuck one paradigm behind.

The uncomfortable truth the field has been “sweeping under the rug,” per the company’s technical post, is that robust policies for moderately complex tasks require tens to hundreds of hours of post-training data collected in deployment conditions. Worse, research has shown that when post-training data is sufficiently dense, policies trained from scratch with no pre-training at all can match post-trained foundation models. Which raises the obvious question the S1 program is built to answer: what is the point of pre-training? Skild’s answer is that pre-training exists — perhaps solely — to enable learning from one or a few examples.

How S1 Works

S1 is trained on episodic data where the task is specified exclusively through an in-context video demonstration. Because that demonstration may come from a different scene, viewpoint, or even embodiment, the policy must implicitly learn the demonstrator’s intent, the functional correspondences between what it sees and what it must do, and how to track task progress — all to predict the right actions. The model is built on NVIDIA AI infrastructure, and its training data strategy deliberately spans sources that trade off differently on hardware proximity, diversity, and cost: teleoperation (closest to the robot, scales worst), egocentric video (scales best, largest domain gap), and everything in between. Skild says that for every dollar spent collecting data, it spends three on quality control — every data point entering pre-training is screened for low-level precision, task coherence, and annotation fidelity.

In meta-learning terms, pre-training is the outer loop that teaches the policy how to learn from context; at inference time, the demonstration drives the inner loop without changing any weights. Tasks are specified by video, not language, because pre-training across highly diverse tasks forces the model to learn intent from demonstration — yielding one set of weights that generalizes to unseen tasks.

The Numbers: 66% vs 9%

The headline result comes from a controlled study comparing ICL-style demonstration prompting against the language prompting popularized by conventional vision-language-action (VLA) models. Skild trained both policies on identical data, identical architectures (save the prompt embedding), and identical compute, sweeping pre-training datasets from 1,000 to 100,000 hours.

On tasks drawn from the pre-training distribution, the two approaches start close: at 1k hours, the language-conditioned policy actually leads with 53% success versus 43% for ICL. But as pre-training scales, S1’s ICL pulls ahead, reaching 96% success on seen long-horizon tasks — which the company notes is very high for multi-minute manipulations.

The dramatic separation happens on out-of-distribution tasks. Language prompting improves only slowly with scale, plateauing at 9% success even at 100k hours of pre-training. The ICL model reaches 66% at the same dataset size — and Skild reports the gap widens exponentially as data increases, which it reads as an encouraging signal for ICL scaling laws.

Skild attributes the gap to two factors. A novel action primitive — flipping a pancake — is exactly where language lacks grounding in actions, while a demonstration shows it directly. And language is often too coarse to specify a novel chaining of known primitives; a demonstration spells out the composition explicitly.

One Demo ≈ 380 Teleoperated Episodes

Perhaps the most commercially meaningful number in the release is the demonstration-efficiency calculation. Skild post-trained a conventional VLA on unseen tasks using 1 to 2,000 teleoperated demonstrations, then asked how many it takes to match what S1 does from a single video. The answer: roughly 380 post-training episodes. Collecting 380 long-horizon demonstrations (4-10 minutes each) takes 50-100 hours of teleoperation. Post-training does eventually surpass ICL — reaching 86% success at 2,000 demos — but the deployment-time economics are stark. The company’s internal timeline for the plant-potting task makes it concrete: soil, a pot, a watering can, and a plant arrived at the office at 8:54 PM, most of the elapsed time went to moving furniture and setting up the scene, and the gap from recorded demonstration to autonomous execution on hardware was 11 minutes.

The unseen-task suite S1 was evaluated on — plant potting, pancake cooking, pour-over coffee making, and kit assembly — runs up to ten minutes and dozens of manipulation steps, driven by a single visual demonstration. Skild highlights emergent behaviors including digging into soil to make room for a plant, pressing a coffee filter into a funnel, and flipping a pancake read straight off the prompt.

Robustness, Recovery, and a Hint of Common Sense

S1’s qualitative results read like a checklist of everything generalist policies have historically lacked. Perturbation tests slid objects away from the robot mid-approach, swapped objects, and changed lighting — none shown in the prompt — and S1 completed tasks regardless. When it fails, it retries rather than blindly proceeding, with recovery observed off the shelf even on out-of-distribution tasks like assembling a skateboard wheel.

The model also exhibits what Skild calls common-sense behavior: when the demonstration waters a plant with a watering can but only a cup is available, S1 uses the cup; when the prompt fills a glass with juice that is already nearly full, it just tops it off. Most interesting is demonstration correction — when a human demonstrator drops an egg prematurely and makes a mess, S1 performs the same step with a controlled motion, treating the demonstration as a specification of the goal rather than a trajectory to reproduce.

Under formal distribution-shift testing across five severity levels (pose shifts, object substitutions of matched affordance, and forced arm-switching), the language-prompted VLA degraded up to three times as much as the ICL policy at the harshest level. VLA policies handle small pose variations fine but collapse once novel motions are required; in-context policies can be re-prompted with an execution plan that fits the deployment conditions.

Why It Matters

Skild frames S1 as the missing link in its data flywheel: if every environmental change requires data collection, fine-tuning, and validation cycles, robots can never keep pace with real-world environments. Setting up new tasks in minutes lets deployments bootstrap from live robot interactions that feed back into pre-training. Sequoia’s Alfred Lin called single-prompt execution of long-horizon tasks “a game changer,” and Skild says S1 is already working with commercial partners.

The competitive context is worth noting: concurrent in-context manipulation work from Generalist AI covers tasks that are either short-horizon or already in the pre-training distribution, per Skild’s reading. S1 is, the company claims, the first robotics foundation model to show in-context learning on 10-minute, never-seen tasks. One year ago Skild’s LocoFormer proved the approach for locomotion, adapting to unseen embodiments by accumulating live experience in its prompt; S1 transplants those lessons to manipulation, with the first short-horizon signs of life appearing roughly six months ago.

The caveat is the one Skild itself flags: performance is measured with human intervention used to recover from failures during grading, cumulative per-step success on long-horizon tasks, and 66% — not 96% — is the number that matters for genuinely novel deployments. But if the exponential widening of the ICL-versus-VLA gap holds as pre-training scales, the economics of robot learning may have just flipped: one video is now worth 380 teleoperated episodes, and that ratio is a moving target.