CoBPE tokenizer trims model sequences 30% at matched compute
TL;DR
- CoBPE represents phrases as a base token plus reusable surface modifiers composed in embedding space, shortening sequences by about 30% versus standard BPE.
- At 780M and 1.3B parameter scales under matched training compute, CoBPE improved average downstream task performance by 1.2 points over BPE.
- Authors Yuval Reif, Guy Kaplan and Roy Schwartz frame the result as moving part of what token sequences carry into structured representations.
A new tokenization scheme shortens language-model input sequences by about 30% without giving up accuracy. In a paper posted to arXiv, Yuval Reif, Guy Kaplan and Roy Schwartz introduce CoBPE, which represents phrases using "a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output."
Tested at 780M and 1.3B parameter scales, "CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute," the authors report.
The wider claim is deliberately narrow: "part of what is now expressed through token sequences can instead be modeled through structured representations." The abstract reports only the averaged downstream gain; per-task numbers, behavior above 1.3B parameters, and transfer to non-English or code corpora are not covered in the summary.
Shared on Bluesky by 1 AI expert
Originally reported by arxiv.org
Read the original article →Original headline: More Than Words: Compositional Tokenization for Efficient Language Models