VLM With Visual Imagination Branch Gains 39 Points on Spatial Reasoning Without Extra Parameters
Summary
WM-VLM demonstrates that attaching a lightweight world-model branch to a frozen VLM—so the model generates intermediate visual states before answering—lifts spatial reasoning by up to 39.25 percentage points. Ablations confirm the gain comes from the visual states themselves: corrupt or remove them, and performance collapses. This is concrete evidence that visual imagination is a separable, trainable skill, not a byproduct of scale.
Originally reported by paper
Read the original article →Original headline: VLM With Visual Imagination Branch Gains 39 Points on Spatial Reasoning Without Extra Parameters