nature.com web signal

Bolmo byteifies OLMo, Llama and Qwen at under 1% pretrain cost

TL;DR

  • Byteification converts existing subword LLMs into byte-level models using 49.1B tokens, which the authors say is less than 1% of typical pretraining.
  • Bolmo 7B, built from Olmo 3 7B, posts a +16.5% absolute STEM improvement over the earlier byte-level BLT 7B.
  • The team byteified four models: Bolmo 7B, Bolmo 1B, Bwen 8B (from Qwen3 8B) and Blama 8B (from Llama 3 8B).

A team led by Benjamin Minixhofer reports in Nature that it can convert existing subword language models into byte-level ones using 49.1B tokens, which the authors say is less than 1% of a typical pretraining budget. The retrofitted Bolmo 7B, built from Olmo 3 7B, beats the earlier byte-level BLT 7B by +16.5% absolute on STEM tasks, while staying close to its subword source on standard benchmarks.

The authors call the procedure "byteification." Its core move is architectural: it "restores the expressivity of subword-level LLM boundaries by non-causally predicting boundaries for the prefill and then predicting during decoding whether a boundary occurs and the next byte," the paper states. Two training stages do the work, 9.8B tokens then 39.3B.

Parameter overhead is small. Bolmo 1B ends up with roughly 0.7% fewer parameters than its source; Bolmo 7B carries about 4.5% more, Blama 8B (from Llama 3 8B) 2.7% more, and Bwen 8B (from Qwen3 8B) 1.5% more. On character-level benchmarks, byteified models "substantially surpassed subword counterparts," the paper reports.

The authors say the converted systems retain "practical inference speeds by efficiently processing byte-level information."

Shared on Bluesky by 2 AI experts