pub.sakana.ai web signal

Sakana AI's SAIL lifts VLM robot task success from 25% to 73%

TL;DR

  • SAIL raises a Gemini Robotics-ER 1.5 trajectory generator from 25% single-shot success to 73% average when given a 45-node inference-time search budget.
  • Real-world validation is thin: five of six successful trials on a single block-placement task, with simulated task scores ranging from 45% (marker) to 100% (bowl).
  • The pipeline retrieves similar demonstrations, simulates candidate trajectories, scores completion from the resulting video, and refines through Monte Carlo Tree Search without touching model weights.

A vision-language model that lands a robot manipulation task just 25% of the time on its first attempt reaches 73% when it is allowed to generate, evaluate and revise candidates at inference. That is the headline result of SAIL, a paper from Sakana AI and the University of Tokyo accepted to IROS 2026.

The setup uses Gemini Robotics-ER 1.5 as both the trajectory policy and the evaluator. SAIL retrieves semantically similar trajectories from an archive, generates candidates in-context, simulates them, scores subtask completion from the resulting video, and then feeds step-level feedback back to the VLM through Monte Carlo Tree Search. The authors describe the approach as "test-time scaling through search and refinement, improving success rates with additional inference computation without updating the VLM's weights."

The 73% average is unevenly distributed across the reported tasks. Placing a bowl works 100% of the time; a banana 95%; a pen 80%. A marker task lands at 45% and a drawer task at 50%. Intermediate search budgets read 55% at six nodes, 65% at fifteen, 71% at thirty, before topping out at 73% with forty-five.

Physical hardware validation is much thinner: five out of six successful trials on a single block-placement task. The authors are upfront about why the write-up stays cautious. Execution is open-loop, without visual feedback during the run. Simulation and refinement burn extra compute per trajectory. And one real-world task with six trials is not much of a robotics track record.

Two of the researchers we follow shared the paper the day it went up.

Shared on Bluesky by 2 AI experts