nature.com web signal

Nature paper converts subword LLMs into byte-level models

TL;DR

  • The paper introduces "byteification," a two-stage procedure that retrofits subword-based LLMs into byte-level models with minimal extra training.
  • Byteified Bwen 8B (built from Qwen3 8B) posts a +16.5% absolute gain on STEM tasks over the BLT 7B byte-level baseline.
  • Total conversion training used 49.1B tokens, which the authors note is under 1% of a typical pretraining budget.

Benjamin Minixhofer and co-authors report in Nature that a subword-based language model can be converted into a byte-level one using less than 1% of a typical pretraining budget. They call the procedure "byteification," a two-stage conversion that first distills subword behavior into new byte-level components, then enables end-to-end learning of byte-level representations. They apply it to Olmo 3, OLMo 2, Qwen3 and Llama 3 checkpoints, releasing them as Bolmo 7B, Bolmo 1B, Bwen 8B and Blama 8B. The paper was published 7 October 2026, and two researchers from our tracker flagged it the same day.

The headline result: Bwen 8B shows a "+16.5% absolute improvement in STEM tasks over BLT 7B," the byte-level transformer baseline. Total training cost comes to "49.1B tokens in total." The architecture is a latent tokenizer language model that predicts byte boundaries non-causally, using "1 byte of future context" during prefill, which the authors say resolves the expressivity mismatch with subword tokenization.

The motivating pitch is in the authors' own words. "Subword tokenization obscures fine-grained information, which is problematic, especially for scientific data—such as computer code or biological sequences," they write. They argue that "models operating on raw text encodings can scale competitively while offering advantages in domains requiring fine-grained textual understanding," and position the method as a bridge: "byteification establishes a connection between existing subword-level LLMs and byte-level LLMs," rather than a from-scratch rebuild.

Byteified models kept inference speeds comparable to their subword parents, the abstract reports. Per-task numbers against the original Qwen3 8B and Llama 3 8B on general benchmarks are not included in the public summary, so the main data point for now is the STEM-vs-BLT contrast.

Shared on Bluesky by 2 AI experts