Code-switching curriculum aligns multilingual BabyLM models
TL;DR
- Pretraining on code-switched text aligns cross-lingual representations, particularly across different scripts, and the alignment survives later monolingual training.
- A curriculum from word-level, to sentence-level, to monolingual documents beat non-code-switched baselines on the BabyLM evaluation suite.
- The two 100M-word corpora were built from the English, Dutch, and Chinese BabyBabelLM datasets, with an LLM inserting the code-switches.
A group of researchers took the observation that bilingual children mix languages mid-sentence and tried it on transformer pretraining. In a preprint accepted to the BabyLM Workshop at EMNLP 2026, Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, and Kianté Brantley pretrain small decoder-only transformers on two 100M-word multilingual corpora. One is built by mixing the English, Dutch, and Chinese BabyBabelLM datasets. The other is the same corpus after an LLM inserts word- and sentence-level code-switches.
"Children in multilingual communities often code-switch, using multiple languages in a single utterance," the authors write. "Can we induce cross-lingual alignment in language models by training on code-switched text?"
Their answer is yes, with a wrinkle about ordering. The paper reports that training on code-switched data "aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents." Under a curriculum that walks from word-level code-switching, to sentence-level code-switching, to plain monolingual documents, the code-switched models outperform baselines trained without it on the BabyLM evaluation suite.
The abstract stops there. It gives no per-language scores, no parameter counts for the "small" decoder-only transformers, and no comparison against non-curriculum code-switched training. Two of the researchers we track shared the link.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence