nature.com web signal

Byteification retrofits Olmo and Llama into byte-level LLMs

TL;DR

  • A new Nature paper introduces "byteification," converting existing subword language models into byte-level models using under 1% of a typical pretraining budget.
  • Bolmo 7B, retrofitted from Olmo 3 7B, posted a +16.5% absolute improvement on STEM tasks over the prior byte-level baseline BLT 7B.
  • The same recipe was run on OLMo 2 1B, Qwen3 8B, and Llama 3 8B, with parameter overhead ranging from -0.7% to +4.5%.

A method the authors call "byteification" retrofits an existing subword language model into one that reads raw bytes, using less than 1% of what a normal pretraining run would cost. The paper in Nature reports that Bolmo 7B, converted from Olmo 3 7B, posted "a +16.5% absolute improvement in STEM tasks over BLT 7B," the prior byte-level baseline.

The recipe was run on four bases. Olmo 3 7B became Bolmo 7B (+330M params, +4.5%). OLMo 2 1B became Bolmo 1B (-10M params, -0.7%). Qwen3 8B became Bwen 8B (+120M, +1.5%). Llama 3 8B became Blama 8B (+220M, +2.7%). Total training ran to 49.1B tokens across two stages: subword-to-byte distillation with a frozen global model, then end-to-end fine-tuning.

Why abandon subwords. Subword tokenization, the paper argues, "obscures fine-grained information crucial for scientific data like computer code or biological sequences." The byteified models "substantially surpassed their subword-level counterpart regarding character understanding," outperforming earlier byte-level approaches on the CUTE and EXECUTE benchmarks. The authors also merged post-trained Olmo 3 checkpoints into Bolmo via task arithmetic, without extra training.

Two of the researchers we track posted the paper on publication day.

No per-task latency figures accompany the abstract, and the largest base tested is 8B parameters.

Shared on Bluesky by 2 AI experts