Byteification retrofits Olmo and Qwen3 into byte-level LLMs
TL;DR
- A Nature paper published 7 October 2026 introduces "byteification," a two-stage recipe for converting existing subword LLMs into byte-level models with under 1% of typical pretraining compute.
- Bolmo 7B, retrofitted from Olmo 3, beat the previous byte-level baseline BLT 7B by 16.5 absolute percentage points on STEM tasks.
- The same recipe produced Bwen 8B from Qwen3 and Blama 8B from Llama 3, each adding roughly 330 million parameters, a 4.5% bump over the base model.
Bolmo 7B, a byte-level model retrofitted from Olmo 3, beats the previous byte-level baseline BLT 7B by 16.5 absolute percentage points on STEM tasks, at a training cost of less than 1% of typical pretraining. The method, described in a Nature paper published 7 October 2026, converts an existing subword model into one that reads raw bytes while adding roughly 330 million parameters, a 4.5% bump on top of Olmo 3.
The authors call the technique byteification. It runs in two stages: a first pass trains local encoder and decoder components while freezing the global model (9.8B tokens, roughly 43B bytes), followed by full end-to-end training (39.3B tokens, around 173B bytes). The same recipe produced Bwen 8B from Qwen3 and Blama 8B from Llama 3.
Tokenization is a known tax on fine-grained text: code, biological sequences, non-English scripts. The paper argues that "byte-level LLMs offer substantial promise as a foundation for future language models," citing "computational efficiency by reducing energy and deployment costs, for reducing biases introduced by English-centric subword tokenization and for enabling applications that require fine-grained textual understanding."
One architectural detail does real work: a non-causal boundary predictor with one byte of lookahead, which resolves a mismatch between how LLMs model causality over tokens and how subword tokenizers themselves peek at future context to place boundaries. The vocabulary doubles from 256 to 512 bytes once boundary markers are included.
The abstract claims the models "maintain practical inference speeds" but the summary skips per-language and per-domain accuracy numbers. Two researchers we track in our Who's Who were already circulating the paper by its publication week.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature