SYNTH paper trains LLMs on 10-140x fewer tokens from Wikipedia
TL;DR
- SYNTH is an open synthetic corpus derived from 58,698 Wikipedia articles that merges pre, mid and post-training into a single stage.
- The authors say models trained on SYNTH reach high factual precision with 10-140x fewer training tokens, and beat filtered web data at iso-compute.
- The team is releasing SYNTH plus three model families, from a 56M Monad up to a 13B/1B-active MoE, under a permissive license.
A team of researchers claims to have collapsed pre-training, mid-training, and post-training into a single stage, using a synthetic corpus derived from 58,698 Wikipedia articles.
The dataset is called SYNTH, and the method is described in a paper on arXiv by Pierre-Carl Langlais and nine co-authors. In the authors' words, SYNTH 'collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds.'
Three model families came out of it: a 56M tiny model named Monad, 0.3B-0.6B dense models called Baguettotron, and a 13B / 1B-active mixture-of-experts.
The headline empirical claim is stated flatly: 'At iso-compute, SYNTH outperforms filtered web data.' Models trained on SYNTH, the authors write, 'achieve high factual precision despite 10-140x fewer training tokens,' which they credit to back-translation from grounded Wikipedia passages, with 'memorization targeted by the seed corpus.'
The abstract names no specific benchmarks and lists no commercial baselines. The authors frame their motivation around a gap in the public record: frontier labs have built internal synthetic sets 'to augment their pre-training data mix,' but 'none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including small ones, remains poorly understood.'
Two researchers we track flagged the preprint the same day. The SYNTH dataset and the Baguettotron models are being released under what the authors call a permissive license.
Shared on Bluesky by 2 AI experts
-
After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data e…
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs