Anthropic and Google bulk-buy pre-2022 books to dodge AI slop
TL;DR
- ISBNdb is brokering bulk print-book orders for AI labs of between 1,000 and 1 million books per engagement, under strict NDAs.
- Court filings show Anthropic hired ex-Google Books engineer Tom Harvey to lead a scanning effort that planned to ingest millions of books.
- Judge William Alsup ruled that destroying print copies during digitization is 'clearly transformative' fair use, giving buyers legal cover.
The strange corner of the training-data market this week: AI companies are paying real money for used printed books published before 2022, because that is now the cleanest large corpus of human-written text left. 404 Media reports that ISBNdb, which runs a large book database, has moved into brokering bulk print orders for AI labs of between 1,000 and 1 million books at a time, under strict NDAs.
The pitch is blunt. Print books from the pre-LLM era are, in ISBNdb's own phrasing, 'structurally guaranteed to be free of this contamination' from AI-generated text, and that matters because models trained on model output degrade. Web scrapes have been getting steadily more polluted with generated content, and the worry about model collapse is real enough that buying secondhand dead-tree stock has become a serious procurement line.
Anthropic is the most exposed name. Internal documents from its copyright lawsuit showed it planned to obtain and scan millions of books, hired Tom Harvey, who previously worked on Google Books, to run the effort, bought stock from Better World Books, and contracted Datamation for both destructive and non-destructive scanning. Google, separately, has been sued by book publishers over training Gemini on copyrighted books. Judge William Alsup has already ruled that destroying a print copy during digitization is 'clearly transformative' fair use, which is the legal cover the rest of the industry is now operating under. ISBNdb concedes the presentation problem out loud: 'The optics problem is real. "AI company destroys two million books" is not a headline that generates sympathy.'
The honest caveat is that the most specific numbers here come from ISBNdb's own marketing and from one professional bookseller who told 404 Media he has been moving 'hundreds of books a week' since April, up from roughly twenty. Take the volumes as reported, not settled. What the reporting does not give you is how much money is changing hands per book, whether any of it reaches authors of long-tail backlists, or how cleanly the Alsup ruling extends from Anthropic's specific scanning setup to a broader secondhand pipeline.
The forward-looking bit is who suddenly matters. Secondhand book marketplaces, obscure-backlist publishers, and anyone sitting on provenance-checked pre-2022 text have real leverage right now. Curation, not scale, is the scarce input.
Shared on Bluesky by 10 AI experts (top 5 by trust)
-
companies are sourcing millions of printed books for AI labs that scan them for training data, then destroy them. printed books are more valuable because they're free of AI slop, and AI labs know destroying them is an "…
View on Bluesky → -
I have definitely noticed an uptick in deep backlist sales from the OSU Press list. It's not the perpetual bestsellers but one-offs from decades ago that even I, after seven years at the press, have barely heard of. www.…
View on Bluesky →
Originally reported by 404media.co
Read the original article →Original headline: 404 Media: Anthropic and Google Are Bulk-Buying Pre-2022 Printed Books to Escape AI Slop in Training Data