nature.com web signal

Nature paper retrofits Llama, Qwen to run on raw bytes

TL;DR

  • 'Byteification' converts subword-based LLMs to byte-level models with 49.1 billion tokens of extra training, in two stages.
  • Bolmo 7B posts a +16.5% absolute improvement on STEM tasks over the BLT 7B byte-level baseline, the paper reports.
  • A non-causal boundary predictor uses one byte of future context, closing an expressivity gap earlier byte-level models left open.

A new paper in Nature describes 'byteification,' a two-stage procedure that turns existing subword-based language models into byte-level models with roughly 49.1 billion tokens of extra training. The authors apply it to four open models — Llama 3 8B, Qwen3 8B, Olmo 3 7B and OLMo 2 1B — producing byte-level counterparts they call Blama 8B, Bwen 8B, Bolmo 7B and Bolmo 1B.

The byteified models 'approach source subword model performance across benchmarks,' the paper reports, while Bolmo 7B posts a '+16.5% absolute improvement in STEM tasks' over BLT 7B, the byte-level baseline.

Under the hood is what the authors call a 'latent tokenizer language model' stack: an mLSTM local encoder, a boundary predictor that peeks one byte ahead, a global transformer over patches, and a local decoder. That peek is the point. 'Subword tokenizers use information about future bytes to place token boundaries,' the paper notes, yet earlier byte-level models operated causally and left that signal on the table.

The work spans the University of Washington, University of Cambridge and Meta, with Benjamin Minixhofer listed first and co-authors including Luke Zettlemoyer, Noah A. Smith and Luca Soldaini. They pitch byte-level LLMs as a foundation with 'potential advantages for computational efficiency,' for 'reducing biases introduced by English-centric subword tokenization,' and for applications spanning code and biological sequences. The abstract does not break out per-benchmark tables or inference-speed numbers against the source subword models.

Shared on Bluesky by 2 AI experts