← All posts / Research

Skild AI's S1 Learns New Robot Tasks From a Single Video — No Fine-Tuning Required

Skild AI's S1 robotics foundation model executes 10-minute unseen manipulation tasks from one video demonstration, hitting 66% step success versus 9% for language-prompted VLAs.

Skild AI's S1 Learns New Robot Tasks From a Single Video — No Fine-Tuning Required

Skild AI has unveiled S1, a robotics foundation model that the Pittsburgh-based company describes as its flagship “in-context learner” for manipulation. The claim is strikingly simple to state and notoriously hard to achieve: show the robot a single video demonstration of a task — a task it has never seen during training — and it executes. No fine-tuning, no post-training, no hours of teleoperation data collection. One video, up to ten minutes of autonomous multi-step execution.

Announced on August 25, 2026 and picked up across the robotics press this week, S1 represents the most aggressive public attempt yet to move robot learning past what Skild calls the “BERT era” — the current paradigm where every new task demands fresh data collection and a specialist fine-tuning run, much as every new NLP application once demanded task-specific BERT fine-tuning before GPT-3’s prompting changed the game.

From BERT-era robotics to prompting robots

The framing in Skild’s technical blog post is deliberate. In language modeling, the transition from BERT-style fine-tuning to ChatGPT-style prompting was driven by in-context learning (ICL): the model acquires a novel concept entirely through the prompt, without touching its weights. Robotics, Skild argues, is stuck one paradigm behind. Deep policies can learn complex behaviors, but robustly executing a new task still requires tens to hundreds of hours of post-training data collected in deployment conditions.

Worse, research has shown that when post-training data is sufficiently dense, policies trained from scratch can match post-trained foundation models — which raises the awkward question of what pre-training is actually for. Skild’s answer: pre-training exists to enable in-context learning for robots, exactly as it did for language.

There is also a practical argument about task specification. For atomic actions, language works fine — “hand me the mug” needs no elaboration. But nobody learns to fold a fitted sheet or whisk egg whites to stiff peaks from a sentence. For delicate, long-horizon work, humans stop describing and start showing. S1 does the same: the task enters the model’s context window as a video, and the policy translates it to its own embodiment and current scene.

What S1 actually demonstrated

The four showcase tasks were entirely absent from S1’s pre-training distribution:

  • Plant potting — digging into soil to make room, seating the plant, handling a watering can
  • Pancake cooking — mixing batter and flipping pancakes in a skillet
  • Pour-over coffee — pressing a paper filter into a funnel, dosing grounds, pouring water
  • Mechanical kit assembly — including assembling a skateboard wheel

Several underlying primitives were missing from pre-training entirely — flipping a pancake, digging into soil, pressing a filter. That makes these tasks more than a recombination of known skills, which is the strongest form of the out-of-distribution claim.

The deployment timeline Skild published for the plant-potting task is the pitch in miniature: props arrived at the office at 8:54 PM, the scene was set by 9:16 PM, one egocentric human video was recorded at 9:22 PM, and the robot was executing the task autonomously by 9:27 PM. Eleven minutes from demonstration to deployment, most of it spent moving furniture.

The numbers: 66% vs 9% on unseen tasks

The core result comes from a controlled scaling study. Skild trained both an ICL policy and a conventional language-prompted vision-language-action (VLA) model on identical data, architectures (aside from the prompt embedding), and compute, across pre-training pools from 1,000 to 100,000 hours.

On seen tasks, the story is nuanced. At 1,000 hours, the language-conditioned VLA actually led — 53% versus 43% for ICL — because compressed language tokens specify a familiar task cleanly while a high-dimensional video prompt invites overfitting at small scale. But ICL overtakes as data grows, which Skild attributes to the ambiguity of language: an instruction admits many valid executions, while a demonstration pins down one mode.

On unseen tasks, the gap is dramatic. Language prompting crept up to a 9% step-success rate at 100,000 hours of pre-training. The ICL model reached 66% at the same scale — a 7× improvement. Skild credits two effects: novel skills (a language prompt has no grounding for an action primitive the model has never performed, while visual correspondences transfer) and compositionality (language is too coarse to specify an unseen chaining of known primitives; a demonstration spells out the composition directly).

Two caveats worth holding onto: these are average per-step success rates, not end-to-end completion rates, and human intervention was used during rollouts to recover from failures so every step could be graded.

One demo ≈ 380 teleoperation episodes

Skild also quantified the economics of prompting. A single in-context video demonstration delivers performance equivalent to roughly 380 post-training teleoperation episodes on a held-out task. For long-horizon tasks over four minutes, collecting 380 demonstrations means 50–100 hours of teleoperation.

Post-training still wins eventually — the VLA hits 86% success at 2,000 demonstrations — but the cost asymmetry is the point, and Skild notes the approaches are complementary rather than mutually exclusive.

Emergent behavior, graceful degradation

Beyond raw success rates, Skild documented behaviors that emerged without targeted data collection. Sliding objects away mid-execution, swapping objects, changing lighting — none present in the prompt — failed to derail the policy. When S1 misses a step, it re-attempts rather than blindly proceeding. It shows glimmers of common sense: prompted with a watering-can demonstration but given only a cup, it uses the cup. And it corrects flawed demonstrations — when a human demonstrator drops an egg prematurely and makes a mess, S1 performs the same step with a controlled motion, treating the demo as a goal specification rather than a trajectory to mimic.

Robustness was stress-tested along a five-level displacement ladder (L1–L5, from identical placement up to forcing half the actions onto the opposite arm). The language-prompted VLA degraded up to three times more than the ICL policy under the harshest disturbances, tolerating small pose variations but collapsing once novel motions were required.

The catch and the competition

Skild bounds its own claim honestly: short, in-distribution tasks of five to thirty seconds don’t need ICL at all, since frontier VLAs handle those zero-shot from language. The case rests on the hard regime — long, unseen, dexterous tasks where one mistake cascades.

The release also sharpens a philosophical split in physical AI. Generalist AI’s GEN-1.5 pursues one-shot adaptation but only across short 3–12 second horizons; Dyna-2 scales million-hour video co-training into generative world models that dream future states. Skild is betting that non-parametric, in-context adaptation can bypass fine-tuning entirely for real deployments. S1 is reportedly already at work with commercial partners, and its minutes-long deployment loop feeds experience back into pre-training — a data flywheel that language-model operators will find familiar.

Whether visual prompting alone can hit the reliability industrial deployment demands remains an open question. But the frontier of zero-gradient robot flexibility just got measurably wider.