Byteification retrofits Olmo and Llama into byte-level LLMs
TL;DR
- A new Nature paper introduces "byteification," converting existing subword language models into byte-level models using under 1% of a typical pretraining budget.
- Bolmo 7B, retrofitted from Olmo 3 7B, posted a +16.5% absolute improvement on STEM tasks over the prior byte-level baseline BLT 7B.
- The same recipe was run on OLMo 2 1B, Qwen3 8B, and Llama 3 8B, with parameter overhead ranging from -0.7% to +4.5%.
A method the authors call "byteification" retrofits an existing subword language model into one that reads raw bytes, using less than 1% of what a normal pretraining run would cost. The paper in Nature reports that Bolmo 7B, converted from Olmo 3 7B, posted "a +16.5% absolute improvement in STEM tasks over BLT 7B," the prior byte-level baseline.
The recipe was run on four bases. Olmo 3 7B became Bolmo 7B (+330M params, +4.5%). OLMo 2 1B became Bolmo 1B (-10M params, -0.7%). Qwen3 8B became Bwen 8B (+120M, +1.5%). Llama 3 8B became Blama 8B (+220M, +2.7%). Total training ran to 49.1B tokens across two stages: subword-to-byte distillation with a frozen global model, then end-to-end fine-tuning.
Why abandon subwords. Subword tokenization, the paper argues, "obscures fine-grained information crucial for scientific data like computer code or biological sequences." The byteified models "substantially surpassed their subword-level counterpart regarding character understanding," outperforming earlier byte-level approaches on the CUTE and EXECUTE benchmarks. The authors also merged post-trained Olmo 3 checkpoints into Bolmo via task arithmetic, without extra training.
Two of the researchers we track posted the paper on publication day.
No per-task latency figures accompany the abstract, and the largest base tested is 8B parameters.
Shared on Bluesky by 2 AI experts
-
Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 buff…
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Retrofitting language models to operate over bytes - Nature