PKU-Alibaba OmniEcho tests spatial audio in embodied agents
TL;DR
- OmniEchoBench includes 197 real-world audio-visual scenes, 2,972 QA pairs, and 900 first-order-ambisonics navigation samples across 30 indoor environments.
- Baseline Qwen3-Omni scores 18.5% overall accuracy on the benchmark; the paper's own OmniEcho model reaches 28.5%.
- The paper is co-authored by Peking University, Alibaba Group, and Tsinghua University, with several authors' work done during Alibaba internships.
On a new spatial-audio benchmark for embodied agents, Qwen3-Omni scores 18.5% overall accuracy and the paper's own OmniEcho model reaches 28.5%. Both numbers are low.
The OmniEcho paper, a collaboration between Peking University, Alibaba Group, and Tsinghua University, introduces OmniEchoBench: 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics audio recorded across 30 indoor environments. The tasks span source direction recognition, 3D localization, source-motion detection, camera-rotation understanding, cognitive-map bird's-eye-view localization, and sound-guided real-world indoor navigation.
The motivation, from the abstract: "Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents." Earlier audio-visual navigation work leaned on simulated acoustics from SoundSpaces; this dataset instead uses "real-world recorded spatial audio," with 12,900 FOA recordings and 600 bird's-eye-view localization questions among the 2,972 QA pairs.
The training set is largely synthetic: 363,193 examples generated by a controllable rendering pipeline. Several authors did the work during internships at Alibaba Group, per the paper's affiliations. The abstract publishes no per-task accuracy breakdown between Qwen3-Omni and OmniEcho.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: PKU's OmniEcho Ships First Unified Spatial-Audio Benchmark for Embodied Agents Across 197 Real-World Scenes