theverge.com web signal

Microsoft: 24 of 8.2M Copilot Chats Matched Books in NYT Suit

Microsoft OpenAI Copyright ai-business

TL;DR

  • Microsoft's summary judgment memorandum says an expert hired by the publishers found 24 Copilot responses across 8.2 million chat logs matched 30-plus words from books at issue.
  • For 202 of the 212 books in the case, the expert found no matches at all; only 10 titles produced any hits.
  • 59,545 records shared at least 16 words with news content used in training, under 1% of the 8.2 million filtered logs.

Microsoft's Sept 4 summary judgment memorandum, filed in the multidistrict litigation before Judge Sidney Stein in Manhattan, says an expert hired by the publishers found 24 Copilot responses across 8.2 million chat logs that shared 30 or more matching words with the books at issue. For 202 of the 212 books, the expert found nothing at all.

The 8.2 million pool was itself filtered. According to The Verge, Microsoft argues the logs were "deliberately selected to be the most damning ones possible," pre-selected because they contained keywords tied to plaintiffs' websites. Even inside that skewed set, 59,545 records shared at least 16 words with news content the model was trained on, under 1% of the total.

The Center for Investigative Reporting's expert separately identified 51 instances of "substantial match." The Times, CIR, and the Authors Guild declined immediate comment.

Microsoft is asking Judge Stein to end the consolidated case at summary judgment on fair use grounds. The MDL groups claims from the Times, the Authors Guild, and roughly 400 local newspapers, one of dozens of AI copyright fights now in discovery.