technologyreview.com web signal

Llama 3.1 trained on 150,000x more words than a kid hears

TL;DR

  • A pre-teen has heard about 100 million words; Meta's Llama 3.1 trained on 15 trillion tokens, roughly 150,000 times more.
  • The 2024 BabyLM champion, GPT-BERT, beat Meta's Llama 2 70B on some benchmarks despite training on 15,000 times less data.
  • Princeton's Brenden Lake trained a model on 61 hours of SAYCam egocentric video and got object recognition without innate biases.

A pre-teen has heard roughly 100 million words. Meta's Llama 3.1 processed 15 trillion tokens during pretraining, about 150,000 times more. Printing that training data would stack past the International Space Station, MIT Technology Review reports; a child's language diet would reach 20 meters. Two experts in our Who's Who directory shared this piece today.

"Claude has seen the amount of language that an entire city will experience in one generation," Georgetown cognitive scientist Ethan Gotlieb Wilcox told the magazine. Toddlers, by contrast, produce grammatically correct sentences after hearing something like 10 to 30 million words.

"If you train GPT-2 on 30 million words, you get a nonsense generator; you don't get a kid," says Stanford's Michael C. Frank.

The gap has produced BabyLM, an annual competition run by UC San Diego's Alex Warstadt that caps training data at 100 million words, or 10 million for a toddler-scale track. The 2024 champion, GPT-BERT, beat Meta's Llama 2 70B on certain benchmarks despite training on 15,000 times less data. Curriculum learning, simple to complex the way a kid encounters language, underperformed expectations. Adding video to text produced no improvement either.

Other groups are chasing the same puzzle from other angles. Princeton's Brenden Lake trained a model on 61 hours of SAYCam egocentric video and got object recognition and word association without the innate biases researchers had previously thought necessary. UC Berkeley's Alison Gopnik points at active learning: "Children are actively exploring, which means that they're actively choosing their own data." Harvard's Elizabeth Bonawitz points at social cognition: children interpret information differently when they recognize a teacher's intent.

Closing the gap has stakes beyond theory. Warstadt argues sample-efficient training could produce serviceable models for minority languages like Czech, Norwegian, or Sami, where only tens of millions of tokens exist.

Shared on Bluesky by 2 AI experts