Nature paper byteifies LLMs for under 1% of pretraining cost
TL;DR
- A new 'byteification' method converts subword LLMs to byte-level operation using 49.1 billion training tokens, less than 1% of a normal pretraining run.
- Byteified versions of OLMo, Llama 3, and Qwen3 approach or exceed their subword sources, with Bolmo 7B gaining 4.5% relative to Olmo 3 7B on aggregate benchmarks.
- Against the earlier byte-level model BLT 7B, the retrofitted Bolmo 7B gains 16.5 absolute percentage points on STEM tasks.
Converting a subword-trained language model to run on raw bytes takes less than 1% of a normal pretraining run, according to a Nature paper published October 7 by researchers at the University of Washington, University of Cambridge, and Meta AI Research. The method, which the authors call 'byteification,' consumed 49.1 billion tokens across two stages: 9.8 billion to distill the subword model into a byte-level one, and 39.3 billion for end-to-end training.
The team applied it to four open-weight checkpoints. The byteified counterparts, Bolmo 1B, Bolmo 7B, Blama 8B, and Bwen 8B, were derived from OLMo 2 1B, Olmo 3 7B, Llama 3 8B, and Qwen3 8B. On aggregate benchmarks Bolmo 7B scored +4.5% relative to its subword source, Blama 8B gained +2.7%, and Bwen 8B gained +1.5%; the smaller Bolmo 1B lost 0.7%. Against the earlier byte-level model BLT 7B, Bolmo 7B gained 16.5 absolute percentage points on STEM.
The architectural wager rests on how subword tokenizers really work. 'The subword tokens themselves are created by taking future context into account,' the paper notes, even though the model only sees past subwords at inference. That motivated a non-causal boundary predictor that peeks one byte ahead during prefill, with the output vocabulary doubled from 256 to 512 so each byte fuses with a boundary symbol.
Post-training transfers across the boundary too: instruction-tuned subword checkpoints can be merged into byteified models via task arithmetic 'without any extra training cost.' Subword models, the paper reports, hit a softmax bottleneck at 200k to 400k vocabulary entries, while byte-level routing offers 'unbounded increase in efficiency for a smooth drop-off in performance.'
Two researchers on our Who's Who roster posted the paper the day it went up.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature