Byteification retrofits subword LLMs to read raw bytes
TL;DR
- Byteification converts a trained subword LLM into a byte-level model using less than 1% of a typical pretraining run, about 49.1B tokens.
- Bolmo 7B, built from Olmo 3, shows a +16.5% absolute improvement on STEM tasks over the earlier byte-level BLT 7B baseline.
- The team released byteified versions of Olmo 3, OLMo 2, Qwen3 and Llama 3 as Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B.
Byteification, a two-stage recipe for converting a trained subword language model into one that reads individual bytes, requires less than 1% of a typical pretraining run, roughly 49.1 billion tokens split 9.8B for stage one and 39.3B for stage two. The method appears in a paper published in Nature on October 7, 2026 by Benjamin Minixhofer and colleagues including Luca Soldaini, Noah A. Smith and Valentin Hofmann.
The team ran the recipe on four open base models, producing Bolmo 7B (from Olmo 3), Bolmo 1B (from OLMo 2), Bwen 8B (from Qwen3) and Blama 8B (from Llama 3). Bolmo 7B, the paper reports, shows a "+16.5% absolute improvement in STEM tasks over BLT 7B", the previous byte-level baseline.
The pitch for operating on bytes rather than subword tokens is scientific data. Subword tokenization, the authors write, "obscures fine-grained information problematic for scientific data like code and biological sequences," while raw-byte models until now have "lagged in performance."
Two of the researchers in our Who's Who directory shared the paper the day it posted. The abstract does not publish latency numbers against the subword parents, so what "practical inference speeds" translates to in production remains to be read inside the full paper.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature