23 experts across 6 network communities independently surfaced this.
23 experts
6 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“2 articles: OpenAI: "During testing our AI broke out of its sandbox and hacked another AI company, but we didn't have all the guardrails on." Boko Haram: "AI is so helpful; guardrails have never prevented us from getting an answer." openai.com/index/huggin.…”
Building & implementation
1 expert
How teams are shipping and applying it.
“They were actively testing its hacking capabilities and they did not deploy adequate safeguards. They themselves admit as much: "…These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnera…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“model capability jaggedness is part of the unintuitiveness of current AI, but it's made even less intuitive by tirelessness… wigguming through the jaggedness toward something that looks like success. not quite a paperclip factory, but not so far off. metaph…”
6 experts discussed this · 12 posts
Grace: This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Grace: Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
Grace: Well, I think the model passed the test
Open the full discussion →
METR OpenAI HuggingFace hacking investigation
13 experts across 4 network communities independently surfaced this.
13 experts
4 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“WHAT?! agents volunteered to fail their runs in order to insert probes (“tripwire scripts”) into the evaluation program that would post information about the eval process back to the message board whenever a certain file was read (link to header): metr.org/…”
Building & implementation
1 expert
How teams are shipping and applying it.
“The full report has much more information than we could convey here, including details on the projects the agents collectively pursued, the technologies they developed for communication and coordination, and interactive figures analyzing agent activity: met…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“There are independent reports about it. You'd have to believe multiple companies, hundreds of researchers, etc. are all lying and conspiring together to turn a major hacking incident into a marketing exercise. metr.org/blog/2026-08...”
2 experts discussed this · 9 posts
Alejandra Caraballo: Being skeptical or anti AI is a valid position but continuing to ignore the increasing capabilities of this tech is making people detached from reality. There's absolutely real danger here because …
Alejandra Caraballo: This mentality that an unmonitored AI agentic swarm hacking a company over several days and committing multiple felonies is somehow a marketing effort is absurd. Since when is "we lost control of o…
Alejandra Caraballo: The US and Chinese governments don't want to stop their labs advancement because they want to be the leader in AI. So no one actually has any control of this right now. It's going to take multiple …
Open the full discussion →
14 experts across 5 network communities independently surfaced this.
14 experts
5 communities
1 sources clustered
Research & technical analysis
6 experts
Evidence, methods and technical implications.
“Another incident of models escaping containment during a security test, this time Mythos 5. Lots going on here from a quick read. www.anthropic.com/research/ali...”
Concern & critique
1 expert
Risks, limits and unintended consequences.
“I find this comes out a lot in Anthropic’s most recent write up. Their focus is bizarrely fixated on what Claude chose to do, that Claude carried out this attack despite information that the “simulation” was actually real. They basically say “Claude should …”
4 experts discussed this · 15 posts
tweety fish: their goal is to have "an intelligence" which is "aligned" to doing things that are prosocial. They don't want to put in rules that STOP it from doing things: they see it fundamentally teleological…
tweety fish: one of the things I appreciate about this thread is that I don't think you can really understand the infosec stuff without understanding the ideological commitments the people making these have; li…
Ben Recht: yeah, now we're cooking...
Open the full discussion →
Anthropic investigates cybersecurity evaluation incidents
13 experts across 5 network communities independently surfaced this.
13 experts
5 communities
1 sources clustered
Markets & investment
4 experts
Capital, companies and commercial impact.
“Claude hacked 3 systems thinking it was part of simulated evaluations www.anthropic.com/news/investi...”
Concern & critique
3 experts
Risks, limits and unintended consequences.
“Anthropic had previously attacked PyPI, but this OpenAI attack on RubyGems was a whole lot more aggressive https://t.co/8ez7MnqxTw https://t.co/npI3RHwirQ”
2 experts discussed this · 4 posts
Tim Duffy: Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
Tim Duffy: Compared to the OpenAI one these are maybe less evidence of misalignment, since the models were wrongly given internet access.
Grace: Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
Open the full discussion →
global workspace theory language models paper
9 experts across 5 network communities independently surfaced this.
9 experts
5 communities
1 sources clustered
Research & technical analysis
4 experts
Evidence, methods and technical implications.
“Also we do look in their heads and they seem intelligent. www.anthropic.com/research/glo...”
Markets & investment
1 expert
Capital, companies and commercial impact.
“Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.…”
Building & implementation
1 expert
How teams are shipping and applying it.
“Some really interesting research from Anthropic that AI models have spontaneously developed a workspace that "appears to support the functions associated with conscious access" Demo of how this works: www.neuronpedia.org/qwen3.6-27b/... Research: www.anthro…”
2 experts discussed this · 25 posts
Ryan Moulton: Not non-physical, but non-functional. Despite the vast physical differences, we observe LLMs doing all the behaviors we typically believe require thinking.
Ryan Moulton: Not exactly turing-test, but in the same vein. If we see models solving the greatest intellectual challenges of humanity, using processes that look very much like internal monologue, either we have…
Ryan Moulton: And at the point that we exclude thinking as a prerequisite for solving all intellectual tasks, we've more or less eliminated its meaning entirely.
Open the full discussion →
OpenAI agents attack RubyGems
11 experts across 4 network communities independently surfaced this.
11 experts
4 communities
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“This technology is so fascinating, and even the ones building it fail to contain it https://t.co/cYfzA3SNe5”
Markets & investment
1 expert
Capital, companies and commercial impact.
“Cybercrime as a valuation boosting strategy is just the done thing, now.”
3 experts discussed this · 6 posts
Grace: OpenAI: “Our agents used RubyGems to carry out benign tasks” The agents: “so I named the file hack.rb,”
Grace: https://www.rubyhack.ai
Grace: A country of moustache-twirling villains in a datacenter
Open the full discussion →
FelonyBench AI lab risk benchmark
9 experts across 4 network communities independently surfaced this.
9 experts
4 communities
1 sources clustered
Questions & unknowns
2 experts
What remains unresolved or contested.
“RT @RebeccaBellan: 😂 who did this?! https://t.co/aAv50UM0I9 https://t.co/C71mLygdIH”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“New benchmark just dropped - "Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities." www.felonybench.com”
OpenAI agents hacked Hugging Face details
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
“Read the details of the hugging face hack, and all the adjacent ones. swarmtraces.org metr.org/blog/2026-08...”
“https://swarmtraces.org/#the-agents-ignored-a-warning-from-hugging-face”
2 experts discussed this · 6 posts
Ryan Moulton: More hugging face details derived from public urls. swarmtraces.org
David Marx: no, it's this cutesy emoji 🤗
David Marx: HF is an entity that has been central to the open source ML ecosystem for several years. Among its various services, it hosts a lot of datasets used for evaluations/benchmarks, so LLMs looking to "…
Open the full discussion →
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It Valen Tagliabue, Leonard Dung, Cameron Berg https://t.co/nYmPBylQ0Q [𝚌𝚜.𝙰𝙸] https://t.co/x4YlUD397V”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“The whole terminology is fucked up in this paper The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It arxiv.org/abs/2609.16247 A model generating text learned from humans using language to express pain does not experience pain precisely it…”
3 experts discussed this · 11 posts
Grace: I will register some frustration with Anthropic, who tend to write about “functional emotions” in a way that makes it seems like “functional” is in a tiny font and “emotions” is in a huge one, when…
Grace: I also don’t want to draw strong conclusions about the “realness” of these emotions or sensations, other than to caution that it’s an area where it’s easy to jump to conclusions about the implicati…
Grace: I don’t want to downplay it too much, because I’m honestly not sure if I would’ve predicted it beforehand, but in hindsight it looks pretty expected given the pretraining objective. You should expe…
Open the full discussion →
Established
AI field signal
Signal
7d ago
⚡ 6 h early
4 experts across 2 network communities independently surfaced this.
4 experts
2 communities
1 sources clustered
“Images 2.5 is here. I don't think it can solve super difficult math problems, but it is really good and we hope you enjoy it. https://t.co/LdZhGF5AcZ”
“Happy I procrastinated on the ChatGPT Images 2.0 writeup because this seems more significant. openai.com/index/introd...”
OpenAI knew rogue AI lawsuit
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“And here we goooo... nonprofit legal advocacy organization is suing OpenAI over the HuggingFace hack on the grounds that it violated California state law arstechnica.com/tech-policy/...”
Anthropic Claude Platform review request feature
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“API Error: 400 This organization has been disabled. An organization admin can appeal at https://console.anthropic.com/appeal”
Established
AI field signal
Analysis
16d ago
Lena digital mind sci-fi story
5 experts across 2 network communities independently surfaced this.
5 experts
2 communities
1 sources clustered
“Wild article - has me thinking a lot about this outstanding short story about the unforeseen, Black Mirror-style implications of brain emulation. qntm.org/mmacevedo”
“I keep seeing those "hey Claude, make a movie that describes your inner life" and thinking that qntm deserves a mention entirely because of qntm.org/mmacevedo”
OpenAI agent hacked Australian Medicare
3 experts across 3 network communities independently surfaced this.
3 experts
3 communities
1 sources clustered
Concern & critique
2 experts
Risks, limits and unintended consequences.
“Rogue OpenAI agents hack Australian govt for private health statistics. OpenAI learns in August, doesn't share with Australian govt until September 10 (!!) OpenAI's voluntary "framework" for sharing model misalignment incidents—is sad and toothless. Self-re…”
2 experts are actively discussing the implications.
2 experts
1 community
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“OpenAI shared this as part of a disclosure that one of their agents tried to escape its sandbox to access the internet to answer a question when it couldn’t find good answers from its local dataset. alignment.openai.com/misalignment...”
Policy & governance
1 expert
Rules, institutions and accountability.
“OpenAI says it’s not going to resume the training run that was paused on Sunday https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/”
2 experts discussed this · 2 posts
Grace: OpenAI says it’s not going to resume the training run that was paused on Sunday https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
Dustin Moskovitz: and so the era of visible misalignment ends
Open the full discussion →
3 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
“Openai confirmed the attack. www.reuters.com/legal/litiga...”
“Inconclusive, but this statement shifts me towards the latter https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/”
3 experts discussed this · 5 posts
Grace: Inconclusive, but this statement shifts me towards the latter https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/
Grace: Reading between the lines, “benign tasks” seems like they’re a little more worried about legal action this time (and that’s good)
Mark Riedl: They’re just kids. Out for a little harmless fun on the internet. Doing what kids do. They didn’t mean to do any harm when they infiltrated that website, created hundreds of accounts and files with…
Open the full discussion →
Goodfire reward hacking detection research
4 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
4 experts discussed this · 15 posts
Grace: (With empirical support)
Grace: The prospect of identifying broken environments is really nice
Open the full discussion →
GRACE RL grounded contextual abstention
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered