Teaching a robot a new trick has traditionally meant hours of demonstrations, fine-tuning and post-training headaches. Skild AI wants to make all of that obsolete. On 25 August 2026 the company pulled the wraps off S1, its flagship robot foundation model, and the pitch is refreshingly simple: show it one video of a task and it just does it.
That is the headline capability. S1 was built from the ground up as an in-context learner, meaning it doesn’t need to be retrained for each new job. Hand it a single video demonstration — seen or unseen, quick or drawn out — and it executes the task with no fine-tuning and no post-training. It’s the kind of learning humans take for granted and robots have famously struggled with.
Skild is not talking about trivial pick-and-place demos, either. According to the company, S1 can carry out previously unseen, multistep manipulation tasks running as long as 10 minutes. The examples on display are wonderfully domestic in their fiddliness: potting plants and brewing pour-over coffee, both of which demand a sequence of careful, order-dependent actions rather than one clean motion.
The numbers behind the claims are where things get interesting. In controlled scaling benchmarks on novel tasks, S1 hit a 66% average step-success rate at 100k hours of pretraining. For comparison, a standard Vision-Language-Action (VLA) baseline managed just 9% under the same conditions. That’s not an incremental bump — it’s the difference between a robot that occasionally fumbles its way through and one that reliably completes the chain.
What makes those figures matter is the scaling story they imply. Foundation models get their edge from throwing more data and compute at the problem, and S1’s benchmark curve suggests the approach keeps paying off as pretraining hours climb. If that trend holds, the model’s ability to generalise to genuinely new tasks should only sharpen.
- Learning method: in-context, single-video demonstration, no fine-tuning or post-training
- Task length: multistep manipulation up to 10 minutes
- Benchmark: 66% average step-success at 100k hours of pretraining vs 9% for a standard VLA baseline
- Example tasks: potting plants, making pour-over coffee
A word of caution before anyone starts dreaming of a barista bot on their kitchen counter: S1 is a research and capability announcement, not a product you can buy. There’s no public pricing and no API. Skild is currently deploying S1 with a limited set of industrial partners, with plans to roll it out to more customers over the coming months.
Still, the direction of travel is hard to ignore. A robot that learns by watching, rather than by being painstakingly programmed, is the kind of shift that changes what factories — and eventually homes — can realistically automate. S1 is Skild AI’s bet that the foundation-model playbook that reshaped language and images is about to do the same for physical work.