Found first: a primary source the press has not covered yet.
By August 2026, 31.1% of tokens in a FineWeb-quality-filtered web crawl were labeled AI-generated, up from 27.5% in June. Researchers from Pangram Labs and the University of Maryland trained 800 language models to measure the pretraining cost, and found that reaching the same loss as a human-only dataset at 20 tokens per parameter requires 1.6× the compute. The paper is available on arXiv.
What the source says
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, and Bradley Emi, working across Pangram Labs and the University of Maryland, applied the Pangram AI detector to FineWeb-filtered web crawls at two points: June 2026 (27.5% AI-generated tokens) and August 2026 (31.1%). They pretrained 800 language models ranging from 19.9M to 973M parameters, varying the ratio of AI to human tokens in each training mix; the scaling law was fitted on models up to 268M parameters and validated on larger held-out sizes. Adding AI tokens initially lowers loss on human text, but the benefit saturates and reverses under data-constrained conditions, and at the August 2026 share of 31.1%, matching human-only training at 20 tokens per parameter requires 1.6× the compute. Hoffmann et al. (2022) Chinchilla scaling laws fail to predict this behavior. The paper proposes a two-term replacement with separate benefit and harm components that reduces to Chinchilla when no AI text is present, cutting prediction error by 41% on held-out model sizes.
Why it matters
Any lab pulling pretraining data from web crawls now faces a quantified compute tax, and the AI token share has been rising month over month with no sign of leveling. Using standard Chinchilla estimates on today's web will systematically underestimate compute requirements and overestimate model quality on human text distributions. The researchers released the WildAI corpus, 83 billion tokens labeled for AI origin, topic, and format, along with all 800 trained models and reproducible code, giving outside teams a direct tool to measure the same effect in their own datasets.