arxiv.org web signal

Paper: Encoder-Free MLLMs Predicted to Catch Up at ~10^22 FLOPs

TL;DR

  • A new scaling-laws study compares encoder-free and encoder-based multimodal LLMs and projects the two curves converging around 10^22 training FLOPs.
  • Removing the visual encoder shifts the compute-optimal allocation toward larger models for multimodal training, but leaves the text objective nearly unchanged.
  • Without an encoder, the language model adapts: bidirectional visual-token interactions grow more useful, visual processing shifts earlier, and expert routing concentrates.

A new arxiv preprint puts a number on when the visual encoder inside a multimodal model stops paying its way: around 10^22 training FLOPs.

Lin Chen and eight co-authors compare scaling laws for encoder-free and encoder-based multimodal large language models. On the text objective the two curves basically overlap. On the multimodal objective they don't. The paper reports that encoder-free models 'underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets.'

The trade-off is not free. Removing the encoder 'shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged.' An encoder-free run wants to spend its budget on parameters rather than tokens.

How a plain language model takes over an encoder's job is the paper's other contribution. Three adaptations emerge with scale: 'bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated.'

The 10^22 figure is a projection, not a completed run. The authors frame encoder-free as 'a promising direction for multimodal pretraining' rather than a settled winner. It lands in a busy stretch: seventy multimodal papers have moved through our tracker in the last ninety days.