paper web signal

CARE Paper Certifies 9.0–10.8× Speedup for VLA Robot Inference

TL;DR

  • CARE certifies 9.0–10.8× inference speedups on four LIBERO suites with OpenVLA-OFT, guaranteeing at 95% confidence that at least 85.8% of reference-solved episodes are preserved.
  • Under tight budgets, accelerator selectors without guarantees exceed the user-specified failure budget in up to 75% of trials, according to the paper.
  • The sequential form of CARE uses 78.9% fewer rollouts than exhaustive evaluation and generalizes to flow-step reduction for π₀.₅ and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.

Fast inference methods for vision-language-action robot policies bust their safety budget in up to 75% of trials when deployed without certification, according to a new arXiv paper by Rui Liu, Tong Zheng, Jindong Gu and Zhipeng Wang.

The authors propose CARE, "an approach for certified accelerator selection." It runs paired rollouts from identical initial conditions and tracks the gap: when the reference policy succeeds but the accelerated one fails, that counts as an acceleration-induced failure. The method then "deploys the fastest certified candidate, falling back to the reference if none qualify."

On four LIBERO suites with OpenVLA-OFT, CARE "certifies 9.0–10.8× speedups while guaranteeing (at 95% confidence) that at least 85.8% of reference-solved episodes are preserved." The sequential variant uses "78.9% fewer rollouts than exhaustive evaluation."

The paper's framing of the hidden risk is blunt. "Acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics," the authors write, arguing that "action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes."

Beyond LIBERO, the authors report the method generalizes to flow-step reduction for π₀.₅ and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter. The abstract reports only aggregate figures. There is no per-task breakdown, no wall-clock latency in milliseconds, and no disclosure of which specific acceleration mechanism wins certification in each suite.

Shared on Bluesky by 1 AI expert