Nature paper retrofits Olmo, Llama, Qwen to run on bytes
TL;DR
- A Nature paper introduces 'byteification,' which converts existing subword language models into byte-level ones using less than 1% of a typical pretraining budget.
- The retrofitted Bolmo 7B shows a +16.5% absolute improvement in STEM tasks over BLT 7B, the current byte-level baseline.
- Training ran in two stages, 9.8B then 39.3B tokens, on roughly 172B tokens from Dolma 3 augmented with 75M CUTE-style tokens.
Subword tokenizers can be swapped out of existing frontier language models at under 1% of a typical pretraining budget, according to a Nature paper published October 7. Minixhofer and colleagues call the method byteification, converting Olmo 3, Llama 3 and Qwen3 into byte-level counterparts (Bolmo, Blama, Bwen) through a two-stage procedure: train a local encoder, decoder and boundary predictor around the frozen original model, then fine-tune end to end.
Bolmo 7B registers a '+16.5% absolute improvement in STEM tasks over BLT 7B,' the current byte-level baseline, while the retrofitted models 'substantially surpassed their subword-level counterpart regarding character understanding.' Training ran on roughly 172B tokens from Dolma 3, with 9.8B tokens in the first stage and 39.3B in the second, plus 75M tokens of CUTE-style data.
The authors argue the economics tilt further at the top end. 'Byte-level LLMs do not suffer from the softmax bottleneck,' they write, with subword models becoming 'Pareto-dominated' once vocabularies grow between 200k and 400k tokens. They frame byteification as removing 'a long-standing performance barrier to end-to-end byte-level language modelling.'
Two researchers we track circulated the paper when it dropped. The abstract does not quantify inference wall-clock cost against the subword originals, and the 1% figure is anchored to the Dolma 3 pretraining pipeline; teams building from a different corpus will not inherit the same economics by default.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature