paper web signal

PKU-Alibaba OmniEcho tests spatial audio in embodied agents

TL;DR

  • OmniEchoBench includes 197 real-world audio-visual scenes, 2,972 QA pairs, and 900 first-order-ambisonics navigation samples across 30 indoor environments.
  • Baseline Qwen3-Omni scores 18.5% overall accuracy on the benchmark; the paper's own OmniEcho model reaches 28.5%.
  • The paper is co-authored by Peking University, Alibaba Group, and Tsinghua University, with several authors' work done during Alibaba internships.

On a new spatial-audio benchmark for embodied agents, Qwen3-Omni scores 18.5% overall accuracy and the paper's own OmniEcho model reaches 28.5%. Both numbers are low.

The OmniEcho paper, a collaboration between Peking University, Alibaba Group, and Tsinghua University, introduces OmniEchoBench: 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics audio recorded across 30 indoor environments. The tasks span source direction recognition, 3D localization, source-motion detection, camera-rotation understanding, cognitive-map bird's-eye-view localization, and sound-guided real-world indoor navigation.

The motivation, from the abstract: "Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents." Earlier audio-visual navigation work leaned on simulated acoustics from SoundSpaces; this dataset instead uses "real-world recorded spatial audio," with 12,900 FOA recordings and 600 bird's-eye-view localization questions among the 2,972 QA pairs.

The training set is largely synthetic: 363,193 examples generated by a controllable rendering pipeline. Several authors did the work during internships at Alibaba Group, per the paper's affiliations. The abstract publishes no per-task accuracy breakdown between Qwen3-Omni and OmniEcho.

Shared on Bluesky by 1 AI expert