Nature paper byteifies Olmo, Qwen, Llama into byte-level LLMs
TL;DR
- A new Nature paper retrofits subword-based LLMs into byte-level models using 49.1 billion tokens of extra training, which the authors call under 1% of typical pretraining.
- The byteified Bolmo 7B reports a +16.5% absolute STEM-task gain over BLT 7B, a byte-level transformer trained from random initialization.
- The two-stage conversion procedure is applied to Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B, producing Bolmo, Bwen and Blama variants.
A new paper in Nature converts existing subword-based language models into byte-level ones using 49.1 billion tokens of extra training. The authors call that 'less than 1% of a typical pretraining budget.'
The method has a name: byteification. A two-stage conversion procedure is applied to Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B, producing retrofits named Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B. 'We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training,' the paper states.
The headline comparison is against BLT, a byte-level transformer built from scratch. 'Bolmo 7B achieved a +16.5% absolute improvement in STEM tasks over BLT 7B, which was trained from random initialization,' the authors report. Retrofitting beat training a byte-level model from zero at a fraction of the compute.
'Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding,' the abstract concludes. The paper's link moved through two researchers in our Who's Who shortly after it posted.
Per-task result tables and inference-cost figures are not in the abstract text as retrieved.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature