paper web signal

ReImaGin lifts multimodal reasoning up to 25% via image gen

TL;DR

  • ReImaGin calls an image generation model mid-reasoning, letting a multimodal LLM issue natural-language commands like 'remove an occlusion' or build a floorplan.
  • Across six visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin beats text-only chains and specialist vision tools by up to 25%.
  • The authors argue fixed-function tools like depth estimators are too rigid; generative image models accept open-ended prompts and return new visual content for reasoning.

ReImaGin, described in a preprint by Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach and Anna Rohrbach, calls a generative image model in the middle of a reasoning chain and posts gains of up to 25% across six visual reasoning benchmarks, per the arXiv listing. The setup lets a multimodal LLM issue natural-language commands to an image generator, with the paper's examples including "removing an occlusion" and generating "a floorplan from multiple disjoint views of a room", instead of routing through fixed specialist tools.

The authors pitch this against the existing menu of vision expert tools such as depth estimation and object detection modules. Those, they write, "remain fundamentally limited by their reliance on narrow, rigid operations that cannot flexibly generate or transform visual content." ReImaGin's alternative accepts open-ended prompts and returns new visual content the reasoning chain can then read from.

The evaluation covers six tasks including multi-view spatial reasoning and collision prediction, with ReImaGin "consistently" beating both text-only reasoning and specialist vision-tool baselines. The abstract publishes no per-task numbers, no model backbones, and no baseline names, only the ceiling figure and the task list.