huggingface.co web signal

OneSearch-VL 8B Beats Qwen3-VL-8B by 20.2 Points on Multi-Image Research

TL;DR

  • OneSearch-VL-8B, built on Qwen3-VL-8B, scores 55.8 on OneSearch-MI-Bench versus 35.6 for the Qwen3-VL-8B tool baseline, a 20.2-point gain.
  • On VideoDR the agent reaches 57.0 against the 30.0 baseline, and lifts a seven-benchmark single-image average from 42.0 to 58.3.
  • The full RL reward combining accuracy, query, trace, and ground signals averages 61.1, 3.8 points above an answer-plus-query baseline.

A new paper on Hugging Face reports an 8B multimodal agent that beats its own base model by 20.2 points on a multi-image research benchmark the same team built. OneSearch-VL-8B, initialized from Qwen3-VL-8B, scores 55.8 on OneSearch-MI-Bench against 35.6 for Qwen3-VL-8B with the same tool access. On OneSearch-Video-Bench the gap is 17.6 points; on VideoDR it widens to 27.0.

The mechanism the authors put at the center is a Visually Grounded Evidence Graph, which they describe as a structure that "links localized visual anchors to real-world entities, records multi-hop relational paths and source-supported facts, and specifies the operations that compose these facts into an answer." From VGEG annotations they derive the Evidence-aware Visual-Grounded Rubric reward, used during RL to score two things explicitly: evidence traceability and visual grounding. The ablation shows the full reward reaches a 61.1 average, versus 57.3 for an answer-plus-query reward, versus 55.8 for the SFT-only baseline.

Training runs in two stages on OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K. Expert trajectories were synthesized with Seed 2.0 Pro; final answers are judged by GPT-4o returning a binary correctness decision, following OpenSearch-VL. The two new benchmarks hold 608 questions in total, and the authors note that "the five compositional categories beyond single-anchor lookup account for 555 questions, or 91.3% of the combined set."

The paper does not list an affiliation in the version served on Hugging Face, and the gains against larger frontier VLMs with tools are not reported. It arrives into a crowded open-source multimodal week on our tracker, alongside KAIST's ME-World and Zhejiang's SpaceCast-Bench. Author Manyuan Zhang and twelve co-authors are listed; affiliation is not.