Nature paper retrofits OLMo and Qwen into byte-level LLMs
TL;DR
- A Nature paper introduces 'byteification,' a two-stage retrofit that converts subword LLMs like OLMo, Qwen and Llama into byte-level models.
- Byteified Bolmo 7B posts a +16.5% absolute improvement on STEM tasks over BLT 7B, with boundary prediction reaching over 99% accuracy.
- The full conversion uses 49.1B training tokens total — 9.8B in stage one, 39.3B in stage two — described as under 1% of a typical pretraining budget.
A new Nature paper introduces 'byteification,' a two-stage procedure that retrofits existing subword-based language models into byte-level ones using 'less than 1% of a typical pretraining budget (49.1B tokens in total).' The authors apply it to OLMo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B, producing byte-level variants they call Bolmo, Bwen and Blama.
Bolmo 7B posts a '+16.5% absolute improvement in STEM tasks over BLT 7B,' the previous publicly available byte-level model of comparable size trained from random initialization. The paper also reports that the boundary predictor — the component that groups bytes into patches for the global transformer — reaches 'over 99% accuracy.'
Subword tokenizers slice text into word fragments, which obscures the individual characters a model needs for code, math or biological sequences. 'Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data—such as computer code or biological sequences,' the authors write. Earlier byte-level attempts handled characters directly but trailed subword models on general tasks.
Byteification splits the conversion in two. The first stage (9.8B tokens) trains the boundary predictor, local encoder, decoder and language-modeling head while the global transformer stays frozen and distills from the source subword model. The second stage (39.3B tokens) unfreezes the stack and trains it end-to-end so the global model learns to use byte-level information directly.
'Our results remove a long-standing performance barrier to end-to-end byte-level language modelling,' the authors write. Two researchers we track shared the paper's link shortly after it posted.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature