Byteification retrofits Qwen, Llama, Olmo to byte-level LLMs
TL;DR
- A Nature paper introduces 'byteification,' a two-stage retrofit that converts existing subword LLMs into byte-level models using less than 1% of typical pretraining budget.
- The byteified Bwen 8B, converted from Qwen3 8B Base, posted a 16.5% absolute improvement on STEM tasks over the BLT 7B byte baseline.
- The method was applied to Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B with roughly 49.1 billion tokens of conversion training.
Research retrofits existing subword language models to operate directly on bytes using less than 1% of typical pretraining compute, according to a paper published October 7 in Nature. The byteified Bwen 8B, built from Qwen3 8B Base, posted a 16.5% absolute improvement over the BLT 7B byte baseline on STEM tasks.
The approach, called "byteification," is a two-stage conversion the authors applied to four model families: Bolmo 7B from Olmo 3 7B, Bolmo 1B from OLMo 2 1B, Bwen 8B from Qwen3 8B Base and Blama 8B from Llama 3 8B. Conversion training ran on roughly 49.1 billion tokens, a fraction of the compute used for the original bases.
Subword tokenization, standard across frontier LLMs, "obscures fine-grained information problematic for scientific data like code or biological sequences," the abstract states. The authors report that byteified models kept most of their subword capability while gaining character-level reasoning strength on tasks such as CUTE and EXECUTE.
"Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding," the authors write. Two researchers we track in our Who's Who directory shared the paper the day it went live.
The architecture, which the paper calls a latent tokenizer language model, uses non-causal boundary prediction that gives the model access to "1 byte of future context" during prefill while keeping decoding causal. Lead author Benjamin Minixhofer is affiliated with the University of Cambridge and Epochs; co-authors are at the University of Washington and the University of Edinburgh.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature