paper web signal

PhysEvo lifts frozen Astra to 84% real-robot success rate

TL;DR

  • PhysEvo reports 84.00% success across 25 trials on five real-world manipulation tasks using an AgileX PiPER arm with a frozen Astra model.
  • On eight tasks the authors flag as challenging for direct Astra, PhysEvo scores 55.00% versus 1.25% for the direct-Astra baseline.
  • A task agent plus a meta-agent that revises its own tools drives the gains, with no model-weight updates or separately trained action policy.

A frozen Astra model, wrapped in a harness that lets a second agent diagnose failures and rewrite its own tools, hits 84.00% success across 25 trials on five real-world manipulation tasks with an AgileX PiPER arm, according to a new preprint on arXiv titled "PhysEvo: Astra Can Act, Let It."

The authors call the setup "physical recursive self-improvement (RSI) around a single frozen model." A task agent runs the robot; a meta-agent reads the resulting trajectories "to diagnose failures, revise tools and skills, and test corrections," and can also improve its own diagnostic tools. No weight updates, no separately trained action policy.

In simulation the paper reports a 68.14/100 average score and 62.00% success across 42 RoboDojo tasks on held-out layouts, versus 47.17% for "RoboDawn's one-shot Astra agent, the strongest published reference in our comparison." On eight tasks the authors flag as challenging for direct Astra, the harness lifts success from 1.25% to 55.00%.

The real-world number is the eye-catching one, but the sample is thin: 25 trials, five tasks, a single AgileX PiPER platform, and the hardest comparison is the authors' own direct-Astra reference rather than an independent frontier VLA. The framing the paper leans on is that PhysEvo "turns the consequences of action into persistent, testable changes to how a frozen model acts and improves."