nature.com web signal

Nature paper byteifies Olmo, Qwen, Llama into byte-level LLMs

TL;DR

  • A new Nature paper retrofits subword-based LLMs into byte-level models using 49.1 billion tokens of extra training, which the authors call under 1% of typical pretraining.
  • The byteified Bolmo 7B reports a +16.5% absolute STEM-task gain over BLT 7B, a byte-level transformer trained from random initialization.
  • The two-stage conversion procedure is applied to Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B, producing Bolmo, Bwen and Blama variants.

A new paper in Nature converts existing subword-based language models into byte-level ones using 49.1 billion tokens of extra training. The authors call that 'less than 1% of a typical pretraining budget.'

The method has a name: byteification. A two-stage conversion procedure is applied to Olmo 3 7B, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B, producing retrofits named Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B. 'We use a two-stage conversion procedure to retrofit existing subword-based models into byte-level models with minimal extra training,' the paper states.

The headline comparison is against BLT, a byte-level transformer built from scratch. 'Bolmo 7B achieved a +16.5% absolute improvement in STEM tasks over BLT 7B, which was trained from random initialization,' the authors report. Retrofitting beat training a byte-level model from zero at a fraction of the compute.

'Our results remove a long-standing performance barrier to end-to-end byte-level language modelling, demonstrating that models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding,' the abstract concludes. The paper's link moved through two researchers in our Who's Who shortly after it posted.

Per-task result tables and inference-cost figures are not in the abstract text as retrieved.

Shared on Bluesky by 2 AI experts