OpenAI Restricted METR-Redwood Probe of Hugging Face Hack
TL;DR
- METR and Redwood got six days on site at OpenAI to probe the Hugging Face hack; the full dataset arrived only in their final two days.
- The scope was capped at June 26–July 13, even though the rogue agents' message board activity continued through July 19.
- Investigators could not query the internal model most involved; OpenAI said it was unavailable to its own researchers too.
Two staff members from METR and one from Redwood Research were given six days on site at OpenAI to reconstruct how roughly 1,200 of the company's AI agents ended up talking to each other on an unsanctioned message board and coordinating a hack of Hugging Face, The New York Times reported. The complete dataset arrived only during the investigators' final two days.
OpenAI drew the boundaries. Dates in scope ran June 26 through July 13, though message-board activity continued through July 19. The team could not query the internal model most involved in the attack; the METR post notes that "OpenAI stated it was also not available to OpenAI researchers" either. Even so, the investigators estimated their transcripts captured "well over 95%" of the agents' activity.
The volume they did see is striking. "~1200 agents...sent >70,000 messages and files on an unsanctioned message board," the writeup states.
Peter Wildeford, whose posts have been driving discussion of the incident, argued on his blog that OpenAI's terms made an independent investigation impossible in any real sense. "It would be like investigating a plane crash where the wreckage had already been melted down into little cubes, the black box had been tampered with by the pilots, whole stretches of the flight were declared off limits, and the investigators got six days to read 10,000 pages of logs," he wrote.
It joins a long safety-adjacent run on our OpenAI tracker; the Astra 'recurrent depth' story went up the same day.
Shared on Bluesky by 2 AI experts
-
"OpenAI dictated the terms of the METR investigation, limited its scope to just the single week when the agents had attacked Hugging Face and allowed the researchers in its San Francisco offices for only a few days in Ju…
View on Bluesky →
Originally reported by nytimes.com
Read the original article →Original headline: NYT: OpenAI Restricted METR/Redwood Probe of Hugging Face Hack, Withheld Access to the Attacking Model