Byteification retrofits subword LLMs to run on raw bytes
TL;DR
- A Nature paper introduces 'byteification,' a two-stage procedure that converts existing subword language models into byte-level ones with minimal extra training.
- Bolmo 7B, a retrofit, posted a +16.5% absolute improvement in STEM tasks over BLT 7B, a byte-level baseline trained from random initialization.
- The full retrofit ran on 49.1B tokens total, 9.8B in stage one and 39.3B in stage two, well under a typical pretraining budget.
A method called byteification, published in Nature, converts existing subword-based language models into byte-level ones with 49.1B tokens of additional training, well below the trillions typically used for pretraining from scratch. Stage one used 9.8B tokens; stage two added 39.3B more.
The authors applied the recipe to produce two retrofits of existing open checkpoints, Bolmo 7B and Bwen 8B. Bolmo 7B "achieved a +16.5% absolute improvement in STEM tasks over BLT 7B," a baseline byte-level model trained from random initialization. Bwen 8B "statistically significantly outperformed all the earlier byte-level models" and "numerically improved in every category."
The motivation is practical. "Subword tokenization obscures fine-grained information, which is problematic," the authors write, flagging scientific data "where meaning depends on the individual characters or bytes." Byte-level models have long been attractive for exactly this reason and prohibitively costly to train from scratch.
The paper hedges on where the gains come from. The authors note that character understanding may be "acquired primarily through scale" and that the Olmo 3 source model was "probably trained on substantially more tokens than the other models" it was compared against. That qualifier limits how cleanly byteification itself can be credited for the headline numbers. The paper was already moving through the researcher circles on our radar by the time of posting, with two tracked analysts sharing it.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature