nature.com web signal

Nature: retrofit converts LLMs to byte-level with <1% budget

TL;DR

  • A two-stage 'byteification' retrofit used less than 1% of a typical pretraining budget, totaling 49.1 billion tokens across both stages.
  • Bolmo 7B, retrofitted from OLMo 3, scored a +16.5% absolute improvement in STEM tasks over the earlier byte-level BLT 7B baseline.
  • The method also converted Qwen3 8B into Bwen 8B and Llama 3 8B into Blama 8B, adding between 1.5% and 4.5% parameters.

Byte-level language models have long trailed subword-based ones in raw performance. In Nature on October 7, a team including Benjamin Minixhofer, Luke Zettlemoyer and Noah A. Smith argues it has closed that gap by converting existing subword models into byte-level ones with "less than 1% of a typical pretraining budget."

The authors call the procedure "byteification." Applied to OLMo 3, it produced Bolmo 7B, which the paper reports achieved "a +16.5% absolute improvement in STEM tasks over BLT 7B," the prior byte-level baseline. The same two-stage conversion was also applied to Qwen3 8B (yielding Bwen 8B) and Llama 3 8B (yielding Blama 8B). Parameter overhead ranged from roughly 1.5% for Bwen up to 4.5% for Bolmo 7B.

The paper frames the finding as removing "a long-standing performance barrier to end-to-end byte-level language modelling," pitching the approach at scientific data such as "computer code or biological sequences" where "meaning depends on the individual characters or bytes."

The whole conversion ran on 49.1 billion tokens, split 9.8 billion in stage one and 39.3 billion in stage two. Two of the researchers we follow were circulating the paper within hours of its release.

Shared on Bluesky by 2 AI experts