arXiv paper: AI web tokens flip from help to harm in LLMs
TL;DR
- Pangram labeled 27.5% of June 2026 web tokens as AI-generated, rising to 31.1% by August, after FineWeb quality filtering.
- For high-human-budget runs AI tokens raise loss almost immediately; data-starved runs see an early benefit that saturates and reverses.
- The authors' new scaling law predicts effects on models up to 3.6x larger with 41% lower error than existing laws across all AI ratios.
For large language models trained on healthy budgets of human text, mixing in AI-generated web tokens starts hurting them almost immediately, according to a new scaling-laws paper on arXiv by Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero and Bradley Emi. The team pretrained 800 language models at varying ratios of AI to human tokens, then fit a scaling law that lets a token's value "change sign while also reducing to Chinchilla in the absence of AI text."
Smaller, data-starved runs do see a honeymoon: AI tokens initially lower loss on human text, but "the benefit saturates as more are added and quickly reverses into harm," the authors write. For models trained on high budgets of human text, "AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it."
The reason this is not academic is the share of the web that is now machine-written. After FineWeb quality filtering, the paper reports that 27.5% of tokens from June 2026 web data were "labeled as AI-generated by Pangram, rising to 31.1% by August." That is the pile every frontier lab is scraping from. Two of the AI experts on our Who's Who tracker had already shared the preprint within a day of its September 30 posting.
Existing scaling laws, the authors note, "fail to predict this behavior." Their replacement, fit on smaller models, extrapolates to models up to 3.6x larger with 41% lower error than the best existing law across all AI ratios.
The recommendations are blunt: filter AI text when the target is human text, repeat human text before expanding datasets with AI-generated web text, and report validation loss on human and AI text separately. The authors also release WildAI, an 83B-token corpus with AI, topic and format labels, alongside all 800 models and code.
Shared on Bluesky by 2 AI experts
-
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text https://arxiv.org/abs/2609.40295
View on Bluesky →
Originally reported by arxiv.org
Read the original article →Original headline: How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text