nature.com web signal

'Byteification' converts subword LLMs to byte-level at 1% cost

TL;DR

  • A Nature paper introduces 'byteification,' a two-stage retrofit that converts existing subword LLMs to byte-level models using about 49.1B tokens.
  • Converted models Bolmo 7B/1B, Bwen 8B and Blama 8B stay within 0.7% to 4.5% of their source model's parameter count.
  • Bolmo 7B is reported to show a +16.5% absolute improvement on STEM tasks over BLT 7B at under 1% of pretraining cost.

A new Nature paper claims it can convert an existing subword language model into one that reads raw bytes, for less than 1% of a typical pretraining budget. The method, which the authors call "byteification," uses a two-stage conversion procedure with about 49.1B tokens of additional training.

The paper, published in Nature on October 7, is credited to Benjamin Minixhofer, Luke Zettlemoyer, Noah A. Smith, Luca Soldaini, Valentin Hofmann and co-authors spanning the University of Washington, Cambridge, Meta AI, EPFL, Stuttgart and Edinburgh. They produce four converted models (Bolmo 7B and 1B, Bwen 8B and Blama 8B) that stay within 0.7% to 4.5% of their source model's parameter count.

Bolmo 7B is reported to deliver a +16.5% absolute improvement in STEM tasks over BLT 7B, and the authors write that byteified models "substantially surpassed their subword-level counterpart regarding character understanding." Their headline claim: "Models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding."

The abstract does not publish per-model inference throughput or latency figures against the original subword versions.

Shared on Bluesky by 2 AI experts