nature.com web signal

Byteification retrofits LLMs to bytes for under 1% of pretrain

TL;DR

  • The paper converts existing subword LLMs into byte-level variants: Bolmo 7B, Bolmo 1B, Bwen 8B, Blama 8B, for under 1% of pretraining cost.
  • Bolmo 7B shows a +16.5% absolute improvement in STEM tasks over the BLT 7B byte-level baseline.
  • The two-stage procedure uses 9.8B tokens in stage 1 and 39.3B in stage 2, plus roughly 75M tokens of synthetic character data.

A new Nature paper introduces "byteification," a two-stage procedure the authors use to convert existing subword-tokenized language models into byte-level ones with, by their count, less than 1% of a typical pretraining budget: 49.1 billion tokens in all.

The retrofitted models carry prefix-B names that nod to their base: Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B. On the headline comparison, Bolmo 7B shows a "+16.5% absolute improvement in STEM tasks" over BLT 7B. "Byteification establishes a connection between existing subword-level LLMs and byte-level LLMs," the paper argues; the pitch is that you do not need to pretrain a byte model from scratch to get one.

The compute split: 9.8 billion tokens in stage 1 (around 43 billion bytes), 39.3 billion in stage 2 (around 173 billion bytes), with roughly 75 million tokens of character-level synthetic data layered in.

Why bother leaving subwords behind, in the authors' framing: the move avoids "the softmax bottleneck" that caps subword compression, allowing "unbounded increase in efficiency for a smooth drop-off in performance." Two researchers in our tracker had circulated the paper within hours.

Shared on Bluesky by 2 AI experts