nature.com web signal

Bolmo 7B: retrofitting subword LLMs to operate over bytes

TL;DR

  • A Nature paper introduces 'byteification,' a two-stage retrofit that turns existing subword LLMs into byte-level models with 49.1B tokens of extra training.
  • Bolmo 7B shows a +16.5% absolute improvement on STEM tasks over the BLT 7B byte-level baseline, the authors report.
  • The team releases three retrofitted checkpoints — Bolmo 7B, Bwen 8B, and Blama 8B — spanning multiple source model families.

A paper published 7 October in Nature by Benjamin Minixhofer and co-authors describes a method called 'byteification' for converting existing subword-based language models into byte-level models with minimal extra training.

The authors introduce three retrofitted checkpoints — Bolmo 7B, Bwen 8B, and Blama 8B — spanning multiple source model families. They report that Bolmo 7B achieved a '+16.5% absolute improvement in STEM tasks over BLT 7B,' the main byte-level baseline. The two-stage conversion totals 49.1B tokens of extra training, which the paper frames as less than 1% of a typical pretraining budget.

Byte-level operation matters where tokenizers blur meaning. 'Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data—such as computer code or biological sequences—where meaning depends on the individual characters or bytes,' the abstract states.

The authors frame the result in strong terms, writing that it removes 'a long-standing performance barrier to end-to-end byte-level language modelling.'

Shared on Bluesky by 2 AI experts