nature.com web signal

Nature paper retrofits OLMo, Qwen, Llama into byte-level LLMs

TL;DR

  • A method the authors call "byteification" converts existing subword LLMs into byte-level models using 49.1B total training tokens, under 1% of a typical pretraining run.
  • Byteified Bolmo 7B shows +16.5% absolute improvement in STEM tasks over BLT 7B, and Bwen 8B outperformed all earlier publicly available byte-level LLMs.
  • The retrofit wraps the source transformer in an mLSTM local encoder, a boundary predictor using one byte of future context, and four mLSTM decoder layers.

A new paper in Nature describes "byteification," a procedure for converting existing subword-tokenized language models into byte-level ones using 49.1B tokens in total, which the authors put at less than 1% of typical pretraining.

The authors (Benjamin Minixhofer, Tyler Murray, Tomasz Limisiewicz, Anna Korhonen, Luke Zettlemoyer, Noah A. Smith, Edoardo M. Ponti, Luca Soldaini and Valentin Hofmann) apply the recipe to four open models. OLMo 3 7B becomes Bolmo 7B. OLMo 2 1B becomes Bolmo 1B. Qwen3 8B Base becomes Bwen 8B. Llama 3 8B becomes Blama 8B.

The headline numbers: Bolmo 7B achieved a "+16.5% absolute improvement in STEM tasks over BLT 7B," and Bwen 8B "outperformed all earlier publicly available byte-level LLMs," according to the paper. Byteified models also "substantially surpassed subword counterparts on character understanding tasks." The authors frame the overall result as showing byte-level models can "approach the capabilities of subword-based systems."

Mechanically, the retrofit wraps the source transformer rather than replacing it. A single-layer mLSTM local encoder contextualizes UTF-8 byte embeddings. A non-causal boundary predictor using one byte of future context decides patch edges, with outputs fused into a 512-token vocabulary. The source transformer is retained as the global model. Four mLSTM layers handle local decoding. Training runs in two stages: 9.8B tokens with the global model frozen, then 39.3B tokens end-to-end.

The upside the authors stress is throughput. Models can be "further sped up by training with higher ratios of bytes per patch, which is possible only to a limited extent in subword-level LLMs." They also report byteified models can inherit instruction-following from their subword post-trained siblings via task arithmetic, without further training on the byteified version. Two researchers in our Who's Who feed posted the paper's link the day it published.

Shared on Bluesky by 2 AI experts