Ai2's byteification retrofits subword LLMs to read bytes
TL;DR
- Byteification converts existing subword LLMs into byte-level models with 49.1B tokens of training, under 1% of a normal pretraining run.
- Authors release Bolmo 7B and 1B from OLMo, plus Bwen 8B from Qwen 3 8B and Blama 8B from Llama 3 8B.
- Bolmo 7B shows a +16.5% absolute improvement on STEM over the prior BLT 7B byte-level system.
Converting an existing subword language model into one that reads individual bytes takes less than 1% of a normal pretraining run, or 49.1 billion tokens, according to a paper in Nature by Benjamin Minixhofer and co-authors at the University of Cambridge, University of Edinburgh, and the Allen Institute for AI.
The method, which the authors call byteification, applies a "two-stage conversion procedure" to subword-trained models and produces byte-level versions: Bolmo 7B and Bolmo 1B from OLMo, Bwen 8B from Qwen 3 8B, and Blama 8B from Llama 3 8B. The paper reports that byteified models "substantially surpassed their subword-level counterpart regarding character understanding," the long-running failure mode in which tokenizers obscure the letters inside words, code, and biological sequences.
On STEM tasks, Bolmo 7B shows a "+16.5% absolute improvement" over BLT 7B, an earlier byte-level system. The architecture pairs shallow local encoders and decoders built from mLSTM layers around a deep global transformer, with non-causal boundary prediction that uses one byte of future context to match subword tokenizer behavior.
Two researchers on our tracker list shared the paper after it ran. The largest model demonstrated is 8B. The paper reports no inference-time throughput figures against the subword baselines.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature