huggingface.co web signal

Valeo.ai Where-OPD Lifts MLLM Perception by 3.23 Points

Multimodal Fine-tuning ai-business

TL;DR

  • Where-OPD trains an MLLM student on image and question alone while a frozen teacher gets textual hints about where question-relevant objects sit.
  • On Qwen3.5-4B, average accuracy across seven real-world perception benchmarks rises 3.23 points despite post-training using only procedurally generated synthetic scenes.
  • Vision-OPD, the prior baseline, gains 8.90 points on V* but loses 10.73 on CountQA, suggesting spatial hints generalize better than cropped-view hints.

The teacher sees where the answer is. The student does not.

That is the move in Where-OPD, a paper from Valeo.ai and Sorbonne Université researchers that applies on-policy self-distillation to multimodal large language models. The teacher is given textual guidance identifying which objects in a procedurally generated scene are relevant to the question and where they sit; the student receives only the image and the question. At inference, the student is on its own.

"Our central hypothesis is that distilling behavior induced by spatially grounded guidance can improve how an MLLM identifies and uses relevant visual evidence, and that these improvements can transfer beyond the synthetic scenes used for post-training," the authors write.

The headline number is the transfer. Training uses only synthetic scenes, but average accuracy on seven real-world perception benchmarks (CVBench, V*, HR-Bench 4K, HR-Bench 8K, ZoomBench, MME-RealWorld, and BLINK) rises by 3.23 points for Qwen3.5-4B, with smaller lifts of 1.07 and 1.29 points for Qwen3.5-9B and Qwen3-VL-4B. On task-specific benchmarks, Qwen3.5-4B gains 7.20 points on ChartQA, 10.11 on EvoChart, 2.53 on CountQA, and 1.93 on OCRBench.

The comparison the paper draws against prior work is pointed. Vision-OPD, which gives the teacher cropped views of question-relevant regions, "improves Qwen3.5-4B by 8.90 points on V* and 14.56 points on ZoomBench, yet decreases performance on CountQA by 10.73 points." The authors argue that is because better views help only tasks that benefit from visual zooming; telling the teacher where to look, by contrast, generalizes across tasks that stitch evidence from multiple regions.

The data pipeline is annotation-free. Object identities, attributes and spatial coordinates come automatically from the procedural scenes, so there is no human-grounded dataset and no external higher-capacity teacher model. The paper reports results across three MLLMs and does not disclose compute cost or behavior beyond 9B parameters. The recipe lands alongside a steady run of multimodal research we have been covering over the past three months.