Qwen Team Releases OmniVChat: Native Audio-Video Dialogue Task, Bench and RL Recipe in One Drop
Summary
Alibaba's Qwen team posted OmniVChat, a package covering native audio-visual dialogue where omni models take simultaneous voice-plus-video input and reply in text without intermediate ASR or captioning. The release bundles OmniVChat-Studio, a multi-agent data engine that synthesizes single- and multi-turn dialogues, OmniVChat-Bench across five capability axes with a human-recorded validation subset, and an RL reward design targeting correctness, efficiency and dialogue naturalness. Paper hit Hugging Face's daily list with 30 upvotes shortly after release.
Originally reported by huggingface.co
Read the original article →Original headline: Qwen Team Releases OmniVChat: Native Audio-Video Dialogue Task, Bench and RL Recipe in One Drop