huggingface.co web signal

Huawei VICS Decoder Cuts VLA Safety Cost by Up to 57.5%

Robotics Safety Agents ai-business

TL;DR

  • A Huawei-led team introduces VICS, a training-free reranker that re-scores a frozen vision-language-action policy's candidate moves by how much safe-completion mass they preserve.
  • On the Safety-CHORES benchmark, the generic variant VICS-G lowered mean cumulative safety cost by 1.9% to 57.5% across six PickUp, ObjectNav and Fetch settings.
  • Task success stayed within 2.5 percentage points of plain policy sampling and episode length within 0.82 steps, with no policy retraining required.

A safe move can still be a trap. That is the opening claim of a new Huawei-led paper, posted to Hugging Face, that goes after a failure mode its authors call a safe dead end: an action that passes a local safety check but leaves a frozen vision-language-action policy with no supported route to finishing the task safely.

"A safe action is not necessarily a viable one," the paper opens. "A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion." The team, led by Tu Nguyen at Huawei's Heisenberg Research Center with co-authors from Huawei Noah's Ark Lab, TU Berlin and the UCL Centre for AI, formalizes this as a feasibility–likelihood gap and proposes a decoder, VICS, that reranks candidate actions by how much "feasible-future mass" each one keeps open.

The generic variant, VICS-G, is training-free: it wraps a frozen policy at decode time, with no rollouts and no retraining. On the Safety-CHORES benchmark — 160 PickUp, 200 ObjectNav and 172 Fetch episodes, each in a base and a safety-aligned checkpoint — the authors report VICS-G "lowers mean cumulative safety cost by 1.9%–57.5% across six Safety-CHORES settings while maintaining near-policy success rates and episode lengths within 2.5 percentage points and 0.82 steps." A symbolic-support variant, VICS-S, gets the largest single joint result they publish: +3.5pp success and 46.0% lower cost on base Fetch.

The listed limitations are blunt. Three simulated task families, no physical robot, execution feedback tested at only one perturbation severity, and theoretical recovery guarantees that hold only when the frozen policy actually covers viable candidates. It joins a steady stream of VLA work on our radar this month, including AE-VLA's dual-arm generalization result from Monday.