AE-VLA lifts dual-arm VLA generalization to 21.53% in sim
TL;DR
- ACG-Bench introduces 23 task-condition pairs across 8 task families, with 6 in-domain conditions and 17 unseen dual-arm compositions covering reorder, sync, and cross-task combinations.
- AE-VLA, combining skill-specific LoRA adapters and arm-wise attention on a π_0.5 backbone, reaches 21.53% simulation success on unseen compositions versus 2.94% for a single π_0.5.
- On physical SO101 robots, AE-VLA averages 39.00% across five unseen conditions versus 10.00% for the strongest baseline, but its in-domain score falls below all three baselines.
A combination of skill-specific LoRA adapters and arm-wise attention lifts a shared dual-arm policy's success on unseen skill compositions from 2.94% to 21.53% in simulation, according to a paper posted to Hugging Face on October 6. The method, AE-VLA, is built on the π_0.5 vision-language-action backbone and evaluated against ACG-Bench, a benchmark the same authors introduce.
ACG-Bench contains 23 task-condition pairs across 8 task families, with 6 in-domain conditions and 17 unseen compositions covering reorder, synchronization, their combination, and cross-task cases. 'The combined AE-VLA setting reaches 21.53%, a gain of 18.59 points over Single and 16.00 over Dual,' the paper reports. Two independently controlled π_0.5 policies — a stronger baseline than one shared model — reached 5.53% on the generalization set; MA-VLA's arm-shuffle augmentation reached 3.06%.
On physical hardware, two 6-DoF SO101 arms running AE-VLA hit 39.00% mean success across five unseen conditions, versus 10.00% for the strongest baseline across 20 trials per condition. The lead is uneven: AE-VLA scores 12/20 to 0/20 on Stack Bowls/Sync but trails 2/20 to 6/20 on Push Cubes/Sync. In-domain, AE-VLA reaches 27.17% in simulation and 60.00% on hardware — below Dual π_0.5's 70.00%. 'The real-robot results thus also show improved generalization with some loss of in-domain success,' the paper states.
The authors call a high-level planner that decomposes novel tasks into familiar skills 'a key next step.' The paper is one of several VLA-robotics results we've tracked this week, alongside work on faster world-action models and longer action chunks.
Originally reported by huggingface.co
Read the original article →Original headline: ACG-Bench Paper Debuts AE-VLA Dual-Arm Policy, Lifts Simulation Success to 21.53% From 2.94%