MM-DiT text tokens hold image semantics, even with no prompt
TL;DR
- The authors train a lightweight bottleneck that pipes MM-DiT contextual tokens into a frozen LLM so it can answer natural-language questions about the forming image.
- Even with an empty prompt, those tokens remain decodable, accumulating scene-specific information from the evolving visual representation.
- Images whose contextual tokens read more clearly earn higher human-preference scores, and a technique called Contextual Alignment reinforces that effect in training.
"Contextual tokens encode a rich, global representation of the emerging scene." That is the central claim of a paper posted to arxiv by Omer Dahary, Etai Sella, Hadar Averbuch-Elor, Daniel Cohen-Or, and Or Patashnik, which pokes at the hidden interior of multimodal diffusion transformers.
The method is deliberately plain. The authors train a "lightweight bottleneck network" that maps intermediate contextual tokens into the input space of a frozen large language model, and then ask the LLM questions about the image that is forming. The reader works surprisingly early in denoising, they report, and resolves finer details as generation proceeds.
The oddest finding arrives when the prompt is empty. "This information remains decodable even when the MM-DiT receives an empty prompt," the paper states, "showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself." Images whose contextual tokens read more clearly, the authors add, tend to receive higher human-preference scores.
Building on that, the paper proposes a training technique called Contextual Alignment, which "explicitly reinforces the visual-semantic information encoded in the contextual tokens" and, the authors say, improves generation quality and distributional coverage. The abstract reports no numeric comparisons.
Originally reported by paper
Read the original article →Original headline: MM-DiT Text Tokens Encode Full Image Semantics—Even With an Empty Prompt