Bolmo 7B retrofits subword LLMs to byte-level operation
TL;DR
- A two-stage "byteification" procedure converts existing subword LLMs into byte-level models with 49.1B additional training tokens, per the Nature paper.
- Bolmo 7B posts a "+16.5% absolute improvement in STEM tasks over BLT 7B," the paper reports.
- Byteified models "substantially surpassed their subword-level counterpart regarding character understanding," according to the authors.
The retrofit took 49.1B additional training tokens. Published in Nature on October 7, 2026, the paper describes a two-stage conversion procedure the authors call "byteification," which turns subword-tokenizer language models into byte-level models.
The motivation, as the authors put it: subword tokenization "obscures fine-grained information, which is problematic, especially for scientific data." Their retrofitted 7B model, Bolmo, reports a "+16.5% absolute improvement in STEM tasks over BLT 7B," and the paper says byteified models "substantially surpassed their subword-level counterpart regarding character understanding."
The architecture uses non-causal boundary prediction during prefill. The authors argue byte-level operation sidesteps the softmax computational bottlenecks that constrain vocabulary expansion in subword systems, enabling "unbounded increase in efficiency" with what they describe as manageable tradeoffs.
The abstract reports "practical inference speeds" but gives no wall-clock latency numbers. Two of the researchers in our tracker shared the paper on the day it posted.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature