arxiv.org web signal

Code-switching curriculum aligns multilingual BabyLM models

TL;DR

  • Pretraining on code-switched text aligns cross-lingual representations, particularly across different scripts, and the alignment survives later monolingual training.
  • A curriculum from word-level, to sentence-level, to monolingual documents beat non-code-switched baselines on the BabyLM evaluation suite.
  • The two 100M-word corpora were built from the English, Dutch, and Chinese BabyBabelLM datasets, with an LLM inserting the code-switches.

A group of researchers took the observation that bilingual children mix languages mid-sentence and tried it on transformer pretraining. In a preprint accepted to the BabyLM Workshop at EMNLP 2026, Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, and Kianté Brantley pretrain small decoder-only transformers on two 100M-word multilingual corpora. One is built by mixing the English, Dutch, and Chinese BabyBabelLM datasets. The other is the same corpus after an LLM inserts word- and sentence-level code-switches.

"Children in multilingual communities often code-switch, using multiple languages in a single utterance," the authors write. "Can we induce cross-lingual alignment in language models by training on code-switched text?"

Their answer is yes, with a wrinkle about ordering. The paper reports that training on code-switched data "aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents." Under a curriculum that walks from word-level code-switching, to sentence-level code-switching, to plain monolingual documents, the code-switched models outperform baselines trained without it on the BabyLM evaluation suite.

The abstract stops there. It gives no per-language scores, no parameter counts for the "small" decoder-only transformers, and no comparison against non-curriculum code-switched training. Two of the researchers we track shared the link.

Shared on Bluesky by 2 AI experts