Byteification retrofits OLMo, Qwen and Llama to byte-level
TL;DR
- Byteification is a two-stage retrofit that converts subword language models into byte-level ones using 49.1B tokens of additional training.
- The retrofitted Bolmo 7B outperforms BLT 7B, a byte-level model trained from random initialization, by 16.5 absolute points on STEM tasks.
- The authors apply the method to OLMo, Qwen3 and Llama 3 bases, releasing Bolmo, Bwen and Blama variants at 1B to 8B parameters.
Byteification, a two-stage retrofit described in a Nature paper published October 7, converts existing subword language models into byte-level ones with 49.1B tokens of additional training. The retrofitted Bolmo 7B, derived from OLMo 3 7B, posts a +16.5 absolute-point improvement on STEM tasks over BLT 7B, a byte-level model trained from random initialization.
"We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training," the authors write. The first stage, per the paper, "aims to quickly learn weights for the local encoder, local decoder, boundary predictor and language-modelling head"; the second stage then updates the whole network "so that it learns to use byte-level information." Of the 49.1B-token budget, 9.8B goes to stage one and 39.3B to stage two.
The authors, who include Benjamin Minixhofer, Tyler Murray, Tomasz Limisiewicz, Anna Korhonen, Luke Zettlemoyer, Noah A. Smith, Edoardo M. Ponti, Luca Soldaini and Valentin Hofmann (with affiliations spanning the University of Washington, University of Cambridge and University of Edinburgh), apply the recipe to several open bases, producing Bolmo 1B from OLMo 2 1B, Bwen 8B from Qwen3 8B Base and Blama 8B from Llama 3 8B. The paper reports that Bwen 8B "statistically significantly outperformed all the earlier byte-level models and numerically improved in every category except GenQA." Two researchers in our Who's Who directory posted the paper the day it dropped.
The authors argue byte-level models carry "potential advantages for computational efficiency" and benefits "for reducing biases introduced by English-centric subword tokenization." Per-benchmark numbers beyond the STEM delta and the GenQA note are not in the abstract.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature