Skild Unveils S1 Robotics Foundation Model That Learns 10-Minute Tasks From One Video, No Fine-Tuning
Summary
Skild AI released S1, a robotics foundation model that executes tasks up to 10 minutes long from a single human video prompt with no fine-tuning. The company reports 66% success on unseen tasks versus 9% for language-prompted VLAs at the same 100k-hour training scale, with demonstrations covering pancake flipping, pour-over coffee, plant potting, and kit assembly. Sequoia's Alfred Lin called single-prompt execution of long-horizon tasks 'a game changer.'
Originally reported by skild.ai
Read the original article →Original headline: Skild Unveils S1 Robotics Foundation Model That Learns 10-Minute Tasks From One Video, No Fine-Tuning