Skild AI teaches robots new tasks from one video
Its S1 robotics model uses video prompting to perform unseen long-horizon tasks without task-specific retraining.
Skild AI says its S1 robotics foundation model can learn a new physical task from a single video demonstration, then execute it on robot hardware without updating its weights or running task-specific post-training.
The company frames the approach as in-context learning for robotics: instead of describing a task in language or collecting hours of teleoperation data, an operator shows the robot what to do. Skild says S1 can compose unfamiliar multistep tasks such as plant potting, pancake making, pour-over coffee and kit assembly, with task horizons up to 10 minutes.
NVIDIA says Skild built S1 on NVIDIA AI infrastructure and uses technologies including Isaac Lab, Isaac Sim, Omniverse and Cosmos across simulation, data generation, training and deployment. The companies are also working on GPU-accelerated simulation solvers for physical contact and manipulation.
The important signal is not that robots are suddenly general household workers. It is that robotics may be getting closer to the prompting dynamic that made language models useful: show a new task at runtime, avoid a fresh data-collection cycle, and adapt to a changed factory or warehouse workflow much faster.
Sources
- NVIDIA Newsroomblogs.nvidia.com
- Skild AI S1 research postskild.ai