Sakana AI's SAIL scales robot imitation via test-time MCTS
TL;DR
- SAIL reframes robot imitation as MCTS over full trajectories, guided by a VLM scorer, step-level feedback, and a retrieval archive of past successes.
- Across six simulation manipulation tasks, average success rose from 25% at one rollout to 73% at 45 MCTS nodes, reaching 95% on HandOverBanana.
- On a real-world BlockIntoBowl task with a LeRobot SO-101 arm, the method succeeded in five of six trials; the paper is accepted to IROS 2026.
Robot imitation learning has a fragility problem: one demonstrated trajectory tends to collapse when the robot faces new starting conditions. A new arXiv paper from researchers at The University of Tokyo and Sakana AI proposes spending test-time compute against that failure mode, running Monte Carlo Tree Search over full trajectories.
The framework, called SAIL, sets things up so "each node is a complete trajectory and edges correspond to trajectory refinements." Guidance comes from three pieces: an archive of past successful trajectories for retrieval, a vision-language model that scores candidates, and step-level feedback that assigns trajectory-aligned scores for iterative refinement. The authors report that "the average success rate across all tasks increases from 25% with a single rollout to 73% with 45 MCTS nodes," and the method hits 95% on the HandOverBanana task.
Six simulation manipulation tasks make up the benchmark: HandOverBanana, HandOverPen, BowlOnRack, DrawerOpen, LaptopClose, and MarkerRemoveLid. For the real-world test, the team ran a BlockIntoBowl task on a LeRobot SO-101 arm and reports the method "succeeded in 5 out of 6 real-world trials." Two of the researchers we track had already flagged the paper by the time it landed in the feed.
The paper is accepted to IROS 2026. The abstract publishes no wall-clock or compute cost figures for the additional MCTS budget.
Shared on Bluesky by 2 AI experts
-
Introducing "Scaling In-Context Imitation Learning" (SAIL) to be presented at #IROS2026. This work is a collaboration between Sakana AI and the University of Tokyo. Blog: pub.sakana.ai/sail Paper: arxiv.org/abs/2603.082…
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: SAIL: Test-Time Scaling for In-Context Imitation Learning with VLM