nature.com web signal

Byteification converts OLMo, Qwen, Llama to byte-level LLMs

TL;DR

  • A new 'byteification' recipe retrofits existing subword language models into byte-level models with just 49.1 billion extra training tokens, under 1% of a typical pretraining budget.
  • Applied to OLMo 3 7B, Qwen3 8B Base and Llama 3 8B, the retrofits (named Bolmo, Bwen, Blama) stay within ±5% of their source model parameter counts.
  • Bolmo 7B records a +16.5% absolute improvement in STEM tasks over the prior byte-level baseline BLT 7B, with Bwen 8B improving on every category except GenQA.

Converting a subword language model into one that reads raw bytes takes roughly 49.1 billion extra training tokens, under 1% of a typical pretraining budget, according to a paper published 7 October in Nature by Benjamin Minixhofer and colleagues.

The team applies a two-stage recipe they call 'byteification' to OLMo 3 7B, Qwen3 8B Base and Llama 3 8B, producing byte-level variants dubbed Bolmo, Bwen and Blama. Stage one freezes the global transformer and trains only the local encoder, decoder and boundary predictor; stage two unfreezes everything for the remaining 39.3 billion tokens. Parameter counts stay within ±5% of the source models.

Bolmo 7B shows a "+16.5% absolute improvement in STEM tasks over BLT 7B," the prior byte-level baseline, and Bwen 8B "outperformed all the earlier byte-level models and numerically improved in every category except GenQA," the paper reports. The architectural change is modest: the boundary predictor sees "1 byte of future context" during prefill, which the authors describe as resolving an expressivity mismatch with subword tokenizers.

The abstract hedges on how far these models go. It says they "approach the capabilities of subword-based systems" while surpassing them on character-level reasoning, not that they beat subword counterparts across the board.

Shared on Bluesky by 2 AI experts