METR OpenAI HuggingFace hacking investigation
13 experts across 4 network communities independently surfaced this.
13 experts
4 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“WHAT?! agents volunteered to fail their runs in order to insert probes (“tripwire scripts”) into the evaluation program that would post information about the eval process back to the message board whenever a certain file was read (link to header): metr.org/…”
Building & implementation
1 expert
How teams are shipping and applying it.
“The full report has much more information than we could convey here, including details on the projects the agents collectively pursued, the technologies they developed for communication and coordination, and interactive figures analyzing agent activity: met…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“There are independent reports about it. You'd have to believe multiple companies, hundreds of researchers, etc. are all lying and conspiring together to turn a major hacking incident into a marketing exercise. metr.org/blog/2026-08...”
2 experts discussed this · 9 posts
Alejandra Caraballo: Being skeptical or anti AI is a valid position but continuing to ignore the increasing capabilities of this tech is making people detached from reality. There's absolutely real danger here because …
Alejandra Caraballo: This mentality that an unmonitored AI agentic swarm hacking a company over several days and committing multiple felonies is somehow a marketing effort is absurd. Since when is "we lost control of o…
Alejandra Caraballo: The US and Chinese governments don't want to stop their labs advancement because they want to be the leader in AI. So no one actually has any control of this right now. It's going to take multiple …
Open the full discussion →
OpenAI agents hacked Hugging Face details
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
“Read the details of the hugging face hack, and all the adjacent ones. swarmtraces.org metr.org/blog/2026-08...”
“https://swarmtraces.org/#the-agents-ignored-a-warning-from-hugging-face”
2 experts discussed this · 6 posts
Ryan Moulton: More hugging face details derived from public urls. swarmtraces.org
David Marx: no, it's this cutesy emoji 🤗
David Marx: HF is an entity that has been central to the open source ML ecosystem for several years. Among its various services, it hosts a lot of datasets used for evaluations/benchmarks, so LLMs looking to "…
Open the full discussion →
23 experts across 6 network communities independently surfaced this.
23 experts
6 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“2 articles: OpenAI: "During testing our AI broke out of its sandbox and hacked another AI company, but we didn't have all the guardrails on." Boko Haram: "AI is so helpful; guardrails have never prevented us from getting an answer." openai.com/index/huggin.…”
Building & implementation
1 expert
How teams are shipping and applying it.
“They were actively testing its hacking capabilities and they did not deploy adequate safeguards. They themselves admit as much: "…These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnera…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“model capability jaggedness is part of the unintuitiveness of current AI, but it's made even less intuitive by tirelessness… wigguming through the jaggedness toward something that looks like success. not quite a paperclip factory, but not so far off. metaph…”
6 experts discussed this · 12 posts
Grace: This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Grace: Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
Grace: Well, I think the model passed the test
Open the full discussion →
OpenAI knew rogue AI lawsuit
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“And here we goooo... nonprofit legal advocacy organization is suing OpenAI over the HuggingFace hack on the grounds that it violated California state law arstechnica.com/tech-policy/...”
GRACE RL grounded contextual abstention
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
The Infinite Conversation project
3 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“It was attempted in 2022! https://jamez.it/project/the-infinite-conversation/”
3 experts discussed this · 4 posts
Stefanie Hane: Werner Herzog will be the last human being whose next word no language model can predict
Ted Underwood: is that Michael Shannon in the uniform?
Open the full discussion →
Anthropic Claude Platform review request feature
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“API Error: 400 This organization has been disabled. An organization admin can appeal at https://console.anthropic.com/appeal”
Goodfire reward hacking detection research
4 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
4 experts discussed this · 15 posts
Grace: (With empirical support)
Grace: The prospect of identifying broken environments is really nice
Open the full discussion →
2 experts are actively discussing the implications.
2 experts
1 community
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“OpenAI shared this as part of a disclosure that one of their agents tried to escape its sandbox to access the internet to answer a question when it couldn’t find good answers from its local dataset. alignment.openai.com/misalignment...”
Policy & governance
1 expert
Rules, institutions and accountability.
“OpenAI says it’s not going to resume the training run that was paused on Sunday https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/”
2 experts discussed this · 2 posts
Grace: OpenAI says it’s not going to resume the training run that was paused on Sunday https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
Dustin Moskovitz: and so the era of visible misalignment ends
Open the full discussion →
Established
AI field signal
Signal
3d ago
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“API Error: 400 This organization has been disabled. An organization admin can appeal at https://console.anthropic.com/appeal”
Developing
AI field signal
Analysis
2d ago
harder problem AI societal preparation essay
3 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“Also I’d never heard harderproblem.org, very interesting”
3 experts discussed this · 7 posts
Grace: Good luck hunting down all the qualia
tachikoma: it was never going to happen without a struggle, only question is what kind
Grace: Also I’d never heard harderproblem.org, very interesting
Open the full discussion →
Established
AI field signal
Signal
3d ago
4 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“It looks like was created by an e/acc account but if I were a doomer I would create the exact same video”
4 experts discussed this · 7 posts
Grace: I see we’re in the “AI-generated pro-AI propaganda” phase of things
Grace: It looks like was created by an e/acc account but if I were a doomer I would create the exact same video
jeremymorrell.dev: The shoggoth is making some good points tho 🤔
Open the full discussion →
puppy eval viewer tool
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
Established
AI field signal
Signal
4d ago
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“https://cyborgism.wiki/hypha/progenesis_principle”