nature.com web signal

'Byteification' retrofits OLMo, Qwen, Llama into byte-level LLMs

TL;DR

  • A Nature paper introduces 'byteification,' a two-stage method to convert existing subword-based language models into byte-level ones using 49.1 billion tokens.
  • The converted versions of OLMo 3, OLMo 2, Qwen3 and Llama 3 add between -0.7% and +4.5% parameters versus their subword originals.
  • Bolmo 7B reports a +16.5% absolute improvement on STEM tasks over the earlier BLT 7B byte-level model.

A new Nature paper describes a method to convert existing subword-based large language models into byte-level models using 49.1 billion training tokens, split into a 9.8 billion-token first stage the authors describe as '<1% of typical pretraining budget' and a 39.3 billion-token second stage. The method is called 'byteification,' and the paper, published October 7, 2026, applies it to four open models: OLMo 3 7B, OLMo 2 1B, Qwen3 8B, and Llama 3 8B. The resulting converted models carry between -0.7% and +4.5% parameter overhead relative to their subword originals.

The authors, led by Benjamin Minixhofer, write that 'models that instead operate directly on the byte encoding of text avoid these limitations, but until now they have lagged behind subword-based models.' Their conversion uses a non-causal boundary predictor with '1 byte of lookahead,' which the paper notes is needed because 'subword tokenizers use information about future bytes to place token boundaries.'

On STEM benchmarks, Bolmo 7B is reported at +16.5% absolute over BLT 7B, an earlier byte-level system. The authors frame the technique as a bridge: 'Byteification establishes a connection between existing subword-level LLMs and byte-level LLMs.'

The team is spread across the University of Edinburgh, the Allen Institute for AI, and the University of Cambridge. No per-task inference-speed numbers appear in the available excerpt, and non-STEM benchmark comparisons are not detailed there.

Two researchers in our tracked set shared the paper on its publication day.

Shared on Bluesky by 2 AI experts