paper web signal

Recursive Harness Distillation lifts robot success 37.3%→64.0%

TL;DR

  • Real-world manipulation success rose from 37.3% to 64.0% using a shared playbook, with no VLA model weights updated.
  • On SimplerEnv Bridge, a light agent with the playbook reached 66.7% success versus 41.7% for the GR00T-only baseline.
  • The same playbook applied to the strong agent pushed its success to 79.2%, above the strong agent running without one.

A robot manipulation method called Recursive Harness Distillation raised real-world task success from 37.3% to 64.0% without updating any model weights, according to a new arXiv paper by Seungyeon Kim, Junhoo Lee, Minkyu Kim, Baekseung Kim and Nojun Kwak.

The setup, in the authors' framing: a strong agent discovers effective interventions by interacting with a vision-language-action (VLA) policy, then distills those interventions into a playbook. A light agent executes with the playbook in hand. Its execution feedback is looped back to sharpen the playbook further.

"A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback," the abstract states. The resulting artifact "enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters."

On SimplerEnv Bridge, the light agent with the playbook reached 66.7% success versus 41.7% for the GR00T-only baseline. The same playbook, applied to the strong agent, pushed that agent to 79.2%. The paper reports the playbook-carrying light agent "outperforms the strong agent without a playbook."

The abstract does not name which VLA models filled the strong and light roles in the real-world runs, nor how many episodes or task categories composed the 37.3%-to-64.0% jump. No independent replication yet.