Paper: Encoder-Free MLLMs Predicted to Catch Up at ~10^22 FLOPs
TL;DR
- A new scaling-laws study compares encoder-free and encoder-based multimodal LLMs and projects the two curves converging around 10^22 training FLOPs.
- Removing the visual encoder shifts the compute-optimal allocation toward larger models for multimodal training, but leaves the text objective nearly unchanged.
- Without an encoder, the language model adapts: bidirectional visual-token interactions grow more useful, visual processing shifts earlier, and expert routing concentrates.
A new arxiv preprint puts a number on when the visual encoder inside a multimodal model stops paying its way: around 10^22 training FLOPs.
Lin Chen and eight co-authors compare scaling laws for encoder-free and encoder-based multimodal large language models. On the text objective the two curves basically overlap. On the multimodal objective they don't. The paper reports that encoder-free models 'underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets.'
The trade-off is not free. Removing the encoder 'shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged.' An encoder-free run wants to spend its budget on parameters rather than tokens.
How a plain language model takes over an encoder's job is the paper's other contribution. Three adaptations emerge with scale: 'bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated.'
The 10^22 figure is a projection, not a completed run. The authors frame encoder-free as 'a promising direction for multimodal pretraining' rather than a settled winner. It lands in a busy stretch: seventy multimodal papers have moved through our tracker in the last ninety days.
Originally reported by arxiv.org
Read the original article →Original headline: Paper: Encoder-Free Multimodal Models Predicted to Catch Encoder-Based at ~10^22 FLOPs