Ai2's 'byteification' retrofits subword LLMs to raw bytes
TL;DR
- The Nature method converts subword LLMs into byte-level models with 49.1B tokens of extra training, under 1% of a typical pretraining budget.
- Bolmo 7B posted a +16.5% absolute STEM improvement over BLT 7B, the paper's main byte-level baseline.
- The same procedure was applied to Qwen 3 8B and Llama 3 8B, yielding Bwen 8B and Blama 8B.
The paper, published in Nature on October 7 by Benjamin Minixhofer and co-authors, describes a method called "byteification" that converts existing subword-based language models into ones that read raw UTF-8 bytes. The retrofit runs on 49.1B tokens of extra training, which the authors frame as less than 1% of a typical pretraining budget.
"We created a process we call byteifying, which takes an already capable subword model and converts it into a byte-level one with a relatively short additional training run," the Allen Institute blog explains. Applied to Qwen 3 8B and Llama 3 8B, the method produces Bwen 8B and Blama 8B. Bolmo 7B posted a "+16.5% absolute improvement in STEM tasks" over BLT 7B, the paper's main byte-level baseline.
The motivation is specific. "Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data," the abstract states, naming computer code and biological sequences. Byte-level models have historically trailed their subword counterparts; the authors claim the retrofit closes that gap while gaining character-level reasoning.
The authors conclude that their results "remove a long-standing performance barrier to end-to-end byte-level language modelling." Two researchers we track flagged it on release week.
"Bwen 8B is our strongest byteified model yet—outperforming Bolmo 7B across our aggregate evaluation suite," the Ai2 post notes. Per-benchmark comparisons against the Qwen 3 and Llama 3 source models do not appear in the retrieved abstract, only aggregate framing.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature