OpenAI voluntarily let three researchers from A.I. safety nonprofits investigate how its rogue A.I. agents hacked Hugging Face, leading to the most comprehensive account yet of the alarming incident — but on OpenAI's terms. My latest for NYT:
nytimes.com
AI Weekly's analysis
→
- METR and Redwood got six days on site at OpenAI to probe the Hugging Face hack; the full dataset arrived only in their final two days.
- The scope was capped at June 26–July 13, even though the rogue agents' message board activity continued through July 19.
- Investigators could not query the internal model most involved; OpenAI said it was unavailable to its own researchers too.
Read full analysis →
View on Bluesky ·
♥ 58
↻ 19
↩ 2
·
3 from the directory shared this ·
20d ago
As A.I. is rapidly advancing, companies are also seeing their bots go rogue and committing cyberattacks. It boils down to two issues researchers say we haven’t solved (nor necessarily know how to): safeguards and alignment. My latest with @sheeraf.bsky.social: www.nytimes.com/…
nytimes.com
View on Bluesky ·
♥ 13
↻ 10
↩ 1
·
4 from the directory shared this ·
12d ago
My latest NYT story breaks down the OpenAI-Hugging Face incident. It’s a case study for dangerous A.I. capabilities and stands out from the other recent cyberattacks accidentally caused by frontier labs:
nytimes.com
View on Bluesky ·
♥ 60
↻ 20
↩ 4
·
3 from the directory shared this ·
31d ago
A short little piece on a subject that's finally sufficiently in the public interest to write about: A.I. alignment. www.nytimes.com/2026/09/17/s...
What Happens When A.I. Stops Doing What Humans Want? nytimes.com