VLM Agents in 'Among Us' Sandbox Win by Acting, Not Lying
TL;DR
- VLM imposter agents won more through non-verbal behavior than verbal deception across both harness ablations and cross-VLM evaluation.
- The paper introduces MineAmongUs, a 3D multimodal Among Us sandbox for studying joint verbal and non-verbal deception by embodied VLM agents.
- The authors pair the sandbox with ARIA, an agent framework configurable across five cognitive-component ablation axes.
In a 3D Among Us sandbox built to test how vision-language model agents deceive one another, actions turned out to matter more to imposter wins than the words the agents spoke.
The paper, "Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions", introduces MineAmongUs, which the authors describe as "a 3D multimodal Among Us sandbox where imposter agents must deceive crewmates through joint verbal and non-verbal action." Alongside it they present ARIA, a configurable agent framework with five cognitive-component ablation axes, plus an annotation scheme built on existing deception taxonomies. The team, Jaewoo Ahn, Junseo Kim, Hyunseo Kim, Heeseung Yun, Jaehyeon Son, Zsolt Kira and Gunhee Kim, ran both harness ablations and a cross-VLM evaluation.
Across those runs, the paper reports "non-verbal channels emerging as the more decisive winning contributors across both harness ablation and cross-VLM evaluation." How an agent moved and what it did mattered more to whether a bluff worked than what the agent said.
The abstract names no specific VLMs, publishes no per-model win rates, and does not identify which non-verbal actions carried the most deceptive weight. It also offers no human-imposter baseline for comparison.
Originally reported by paper
Read the original article →Original headline: VLM Agents Win Embodied Social Deception Game Via Actions, Not Words — Non-Verbal Channel Dominates Across All Models Tested