nature.com web signal

Nature paper retrofits OLMo, Qwen and Llama into byte LLMs

TL;DR

  • Byteification retrofits existing subword LLMs into byte-level ones using 49.1B tokens, less than 1% of a typical pretraining budget.
  • Bolmo 7B records a +16.5% absolute improvement in STEM tasks over BLT 7B, the paper reports.
  • The authors apply the method to OLMo, Qwen3 8B Base and Llama 3 8B, producing Bolmo, Bwen and Blama models.

The paper reports that converting an existing subword language model to operate at the byte level takes "49.1B tokens in total," which the authors call "less than 1% of a typical pretraining budget."

The method is called byteification. Published in Nature on 7 October 2026, Benjamin Minixhofer and co-authors apply it to OLMo, Qwen3 8B Base and Llama 3 8B, producing four byteified models they name Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B.

On their own numbers, Bolmo 7B records "a +16.5% absolute improvement in STEM tasks over BLT 7B."

The stake, as a companion Nature News & Views piece frames it, is character-level reading: most LLMs cannot evaluate text at the letter level because they encode words as tokens. A retrofitted model can count the letter "i"s in "artificial intelligence." Two of the researchers on our Who's Who tracker posted the source link.

The architectural detail from the abstract is a two-stage conversion in which the boundary predictor uses "1 byte of future context" to pick patch boundaries, wrapping the source transformer in an mLSTM local encoder. Per-task latency numbers are absent from the retrieved abstract; "practical inference speeds" is the authors' own phrase.

Shared on Bluesky by 2 AI experts