ISBNdb Pitches Pre-2022 Print Books as Slop-Free AI Training Data
TL;DR
- ISBNdb is brokering bulk purchases of printed books published before 2022, pitching them as training data structurally guaranteed to be free of AI-generated slop.
- The service advertises strict NDAs on every engagement, conceding that 'AI company destroys two million books' is not a sympathetic headline.
- Judge William Alsup already ruled that Anthropic's practice of buying, scanning, and destroying print books was 'clearly transformative' fair use.
The clean training data problem has a physical answer: cut the spines off old paperbacks. 404 Media reports that book-database company ISBNdb is openly brokering bulk purchases of printed books for AI firms, pitching pre-2022 volumes as the last large pool of text that is 'structurally guaranteed to be free of this contamination' from generative AI output.
ISBNdb's own sales copy, quoted in the piece, is that 'the world's best AI training data is sitting on a shelf,' representing 'curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate.' Orders start at 1,000+ units, with a 'strict NDA on every engagement,' because ISBNdb concedes the optics: 'AI company destroys two million books' is not a headline that generates sympathy. Booksellers on Alibris and Biblio told 404 they noticed the change starting around April, describing 'not just the quantity, but the weirdness of the orders.'
The legal ceiling on all of this is Judge William Alsup's ruling in the Anthropic case, which found that buying printed books, cutting the spines for faster scanning, and destroying the paper originals was 'clearly transformative' fair use. That decision gives the whole practice cover: buy the physical copy, digitize it, dispose of it, and treat the purchase receipt as the license. Anthropic's own scanning was handled by a firm called Datamation, and its stock came from marketplaces including Better World Books.
The honest caveat is that most of this reporting rests on ISBNdb's own marketing copy and on anonymous booksellers describing a spike. 404 does not name the AI buyers currently placing the bulk orders, and there is no clean number for how many titles have already gone through the shredder. Alsup's ruling is one district decision, and a different court could yet reopen the fair-use question.
What is worth watching is the moat this creates. If the pre-2022 print corpus is genuinely finite and increasingly locked up under NDAs, the labs that hoover it up first walk away with training data no competitor can rebuild from a web crawl, and used-book aggregators quietly become an AI-supply-chain node in a way nobody priced in a year ago.
Shared on Bluesky by 10 AI experts (top 5 by trust)
-
companies are sourcing millions of printed books for AI labs that scan them for training data, then destroy them. printed books are more valuable because they're free of AI slop, and AI labs know destroying them is an "…
View on Bluesky → -
I have definitely noticed an uptick in deep backlist sales from the OSU Press list. It's not the perpetual bestsellers but one-offs from decades ago that even I, after seven years at the press, have barely heard of. www.…
View on Bluesky →
Originally reported by 404media.co
Read the original article →Original headline: AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop