Allen AI retrofits OLMo, Qwen3, Llama 3 to byte-level models
TL;DR
- A new method the authors call byteification converts existing subword language models into byte-level ones using 49.1B tokens, less than 1% of typical pretraining.
- The team released byteified variants of three open families: Bolmo 7B and 1B from OLMo, Bwen 8B from Qwen3, and Blama 8B from Llama 3.
- Bolmo 7B posts a 16.5-point absolute gain on STEM tasks over BLT 7B, a prior byte-level model of comparable size.
A method the authors call "byteification" converts existing subword-based language models into byte-level ones for less than 1% of typical pretraining cost, researchers from the Allen Institute for AI, Cambridge, Washington, Edinburgh, Stuttgart and Potsdam report in a Nature paper published October 7. The team released byteified versions of three open families: Bolmo 7B and 1B from OLMo, Bwen 8B from Qwen3, and Blama 8B from Llama 3.
On STEM evaluations, Bolmo 7B posts a "+16.5% absolute improvement in STEM tasks over BLT 7B", a prior byte-level model of comparable size, with the full retrofit consuming "49.1B tokens in total". The authors frame the motivation plainly: "subword tokenization obscures fine-grained information, which is problematic, especially for scientific data". Computer code and biological sequences are the two examples they cite, both domains where individual characters carry meaning the tokenizer tends to smear.
Byte-level architectures have trailed subword ones at scale because training them from scratch is prohibitive. The paper's claim is that you can skip the from-scratch step. Two researchers we follow on our Who's Who list circulated the paper the day it published.
The authors close by arguing byteified models "scale competitively while offering advantages in domains requiring fine-grained textual understanding".
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature