OpenAI Tells Court It Cannot Search Training Data; Product Page Calls This Core Feature
NEW YORK — Responding to a motion for sanctions filed by The New York Times and co-plaintiff publishers, OpenAI informed a federal judge this week that its systems are technically incapable of searching their training data to locate specific copyrighted articles — a limitation, the company said, it had been unable to clarify earlier due to the complexity of its internal architecture.
The filing was prepared using OpenAI's in-house legal research tool.
The Times's motion alleges that OpenAI misrepresented its search capabilities to the court for nearly two years while withholding datasets and ChatGPT conversation logs, and asks for attorney fees to cover the cost of recovering "improperly withheld" evidence. OpenAI said its responses constituted "good faith characterizations" of systems that do not store training text in a format permitting targeted retrieval.
The characterization landed in the same news cycle as a company blog post describing ChatGPT as "the world's most powerful research and analysis tool," capable of finding, summarizing, and reasoning across "virtually unlimited bodies of text."
Asked to reconcile the two positions, an OpenAI representative said they addressed "different technical domains."
"The model does not retrieve training data the way one might query a database," the representative said. "What it does is synthesize information across an enormous corpus with unprecedented accuracy and speed. These are distinct capabilities."
OpenAI has not described how it will demonstrate that evidence it cannot locate does not exist.
"We can tell you what the internet said about anything," a spokesperson said. "We cannot tell you where we got it."