RoboFoundry Frames Robot Stack as One Self-Evolving Policy
TL;DR
- RoboFoundry improves GPT-5.5 by 27.8% on EmbodiedBench and lifts Qwen3.7-Plus to 70.3%, near GPT-5.5's 72.7%.
- It evolves two surfaces: a context/memory system and a hierarchical skill system with failure-conditioned recovery, updating both from execution traces.
- On RoboMemArena for long-horizon memory, it beats every baseline by at least 39.0%, including baselines assisted by external foundation models.
A new framework called RoboFoundry improves GPT-5.5 by 27.8% on EmbodiedBench and brings Qwen3.7-Plus to 70.3%, close to GPT-5.5's 72.7%, by treating an embodied agent's supporting stack — memory, context, skills, action interfaces — as one policy that evolves rather than tuning each part separately. The paper on arXiv calls the approach 'Self-Evolving System-as-Policy.'
Evolution runs over two surfaces. A context system manages active internal context and persistent file-system memory. A hierarchical skill system organizes atomic skills, reusable compositions, and failure-conditioned recovery. Execution traces get diagnosed for capability gaps and then converted into task-specific system updates, with recurring improvements promoted back to the general system.
The framing rests on a specific claim: 'interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes,' the authors write. A shared semantic interface separates 'embodiment-invariant decisions from embodiment-specific execution,' which is what the authors argue lets evolved capabilities move across different robots.
On RoboMemArena for long-horizon memory, RoboFoundry beats every baseline by at least 39.0%, including baselines assisted by external foundation models. On LIBERO-PRO, gains against a Cap-Agent0 reference run from 243.8% to 679.7% depending on the perturbation type — a larger relative number than the EmbodiedBench headline, but against a specific weaker baseline rather than a frontier system. The authors also report zero-shot transfer and online evolution in real-world robot deployments; the abstract names no platforms and offers no third-party replication.
Originally reported by paper
Read the original article →Original headline: RoboFoundry Treats Full Robot Agent Stack as One Evolving Policy, Gains 27.8% on EmbodiedBench