PhysEvo: 84% Real-World Robot Success With a Frozen Astra, No Weight Updates

Found first: a primary source the press has not covered yet.

A paper posted to arXiv on October 6 introduces PhysEvo, a framework that wraps a frozen Astra vision-language model in a dual-agent loop, letting manipulation competence accumulate through environmental trial with no weight updates and no separately trained action policy. On an AgileX PiPER arm across 25 real trials, the system reached 84% task success. The paper is on arXiv.

What the source says

The work is by Wenqing Tian, Zeyu Zhang, Zhaocheng Liu, Fengwei Liu, Qiang Liu, and Liang Wang; affiliations are not listed on the abstract page. PhysEvo runs two agents: a task agent that executes robot tasks, and a meta-agent that reads the resulting trajectories, diagnoses failures, revises tools and skills, validates corrections, and redeploys. In simulation on RoboDojo (42 tasks, held-out layouts), PhysEvo reached 62.00% success and a 68.14/100 average score, against 47.17% for the RoboDawn Astra baseline. On the eight hardest tasks, where Astra alone almost never succeeds, PhysEvo reached 55.00% success versus 1.25% for the baseline. On five real-world tasks (25 trials, AgileX PiPER), the framework reached 84.00% success and a 90.60/100 average score.

Why it matters

The paper frames this as physical recursive self-improvement: a frozen frontier model getting better at manipulation purely by acting in the world, with a meta-agent absorbing the failure signal. Nothing in the loop modifies the model's weights. For practitioners deploying VLMs on hardware, this opens a path to improving real-world robot performance without waiting for a new model release or running a fine-tuning pipeline. The gap on the hardest tasks is the clearest signal: 1.25% to 55.00% without touching the weights. The real-world evaluation is small (5 tasks, 25 trials), so the scope of generalization beyond the tested hardware and task set is still an open question.