arxiv.org web signal

CoBPE tokenizer trims model sequences 30% at matched compute

TL;DR

  • CoBPE represents phrases as a base token plus reusable surface modifiers composed in embedding space, shortening sequences by about 30% versus standard BPE.
  • At 780M and 1.3B parameter scales under matched training compute, CoBPE improved average downstream task performance by 1.2 points over BPE.
  • Authors Yuval Reif, Guy Kaplan and Roy Schwartz frame the result as moving part of what token sequences carry into structured representations.

A new tokenization scheme shortens language-model input sequences by about 30% without giving up accuracy. In a paper posted to arXiv, Yuval Reif, Guy Kaplan and Roy Schwartz introduce CoBPE, which represents phrases using "a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output."

Tested at 780M and 1.3B parameter scales, "CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute," the authors report.

The wider claim is deliberately narrow: "part of what is now expressed through token sequences can instead be modeled through structured representations." The abstract reports only the averaged downstream gain; per-task numbers, behavior above 1.3B parameters, and transfer to non-English or code corpora are not covered in the summary.

Shared on Bluesky by 1 AI expert