ODU-Bench: Gemini 3.1 Pro Recovers Only 44.7% of Intent Cues
TL;DR
- On ODU-Bench, Gemini 3.1 Pro recovered only 44.7% of context-dependent information needed to infer user intent from speech and vision.
- Eleven of the 14 multimodal language models tested logged false-trigger rates above 50% on non-demand scenarios.
- The benchmark, posted September 18, 2026 by an 18-author team led by Qi Chen, spans single-turn and multi-turn interactions across five dimensions.
A new multimodal benchmark called ODU-Bench finds that Gemini 3.1 Pro recovers only 44.7% of context-dependent information when trying to work out whether a user is actually making a request through combined speech and vision.
Eleven of the 14 multimodal language models the authors tested logged false-trigger rates above 50% on non-demand scenarios, treating more than half of casual, request-like remarks as commands. The arxiv preprint, posted September 18, 2026 by an 18-author team led by Qi Chen, argues that existing benchmarks overlook a capability voice assistants have to get right: determining whether a demand exists at all, then inferring intent from dialogue history and surrounding cues.
The authors call out "underspecified requests, disfluent speech, and noisy acoustic environments" as the failure modes their benchmark deliberately targets. ODU-Bench was assembled via taxonomy-driven methods, agentic video generation, human-recorded interactions and annotated verification, and it evaluates single-turn and multi-turn interactions across five dimensions.
By the time our tracker picked the paper up, two researchers on our watchlist had already shared it.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction