nature.com web signal

Bolmo 7B retrofits subword LLMs to byte-level operation

TL;DR

  • A two-stage "byteification" procedure converts existing subword LLMs into byte-level models with 49.1B additional training tokens, per the Nature paper.
  • Bolmo 7B posts a "+16.5% absolute improvement in STEM tasks over BLT 7B," the paper reports.
  • Byteified models "substantially surpassed their subword-level counterpart regarding character understanding," according to the authors.

The retrofit took 49.1B additional training tokens. Published in Nature on October 7, 2026, the paper describes a two-stage conversion procedure the authors call "byteification," which turns subword-tokenizer language models into byte-level models.

The motivation, as the authors put it: subword tokenization "obscures fine-grained information, which is problematic, especially for scientific data." Their retrofitted 7B model, Bolmo, reports a "+16.5% absolute improvement in STEM tasks over BLT 7B," and the paper says byteified models "substantially surpassed their subword-level counterpart regarding character understanding."

The architecture uses non-causal boundary prediction during prefill. The authors argue byte-level operation sidesteps the softmax computational bottlenecks that constrain vocabulary expansion in subword systems, enabling "unbounded increase in efficiency" with what they describe as manageable tradeoffs.

The abstract reports "practical inference speeds" but gives no wall-clock latency numbers. Two of the researchers in our tracker shared the paper on the day it posted.

Shared on Bluesky by 2 AI experts