nature.com web signal

Ai2 byteification retrofits subword LLMs to run on raw bytes

TL;DR

  • Byteification converts existing subword LLMs into byte-level models using under 1% of a typical pretraining budget, 49.1B tokens total.
  • Bolmo 7B shows a +16.5% absolute improvement on STEM tasks over BLT 7B, with gains on the CUTE and EXECUTE character-level benchmarks.
  • The team released byteified variants of Olmo, Qwen3 8B Base and Llama 3 8B as Bolmo 7B/1B, Bwen 8B and Blama 8B.

Less than 1% of a typical pretraining budget was enough to retrofit existing subword-based language models to operate directly on raw bytes. A Nature paper published October 7 by Benjamin Minixhofer and co-authors introduces a method called byteification, which the authors say took 49.1B tokens in total to apply across models including Olmo, Qwen3 8B Base and Llama 3 8B.

The framing is blunt. "Models that instead operate directly on the byte encoding of text avoid these limitations, but until now they have lagged behind subword-based models in performance," the paper reports. Subword tokenization obscures the fine-grained text structure that matters for computer code and biological sequences; byte-level models historically closed that gap, but typically by training from scratch.

The released checkpoints are Bolmo 7B and Bolmo 1B from Olmo, Bwen 8B from Qwen3 8B Base, and Blama 8B from Llama 3 8B. On aggregate benchmarks, byteified Bolmo 7B hits a +16.5% absolute improvement in STEM tasks over BLT 7B, and the authors say byteified models "substantially surpassed" their subword counterparts on the CUTE and EXECUTE character-level benchmarks. Per Ai2's own write-up, Bwen 8B is the strongest of the four on aggregate.

Architecturally, the move is a non-causal boundary predictor with a 1-byte lookahead, which the paper credits with restoring expressivity equivalent to subword tokenization. Training runs in two stages: first the local encoder, decoder, boundary predictor and language-modelling head; then the whole network. By the time the paper landed, two of the researchers we track in our Who's Who had already posted the link.

The abstract publishes no wall-clock inference numbers, only a claim of "practical inference speeds," and the headline STEM figure is one aggregate against one byte-level baseline.

Shared on Bluesky by 2 AI experts