23 experts across 6 network communities independently surfaced this.
23 experts
6 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“2 articles: OpenAI: "During testing our AI broke out of its sandbox and hacked another AI company, but we didn't have all the guardrails on." Boko Haram: "AI is so helpful; guardrails have never prevented us from getting an answer." openai.com/index/huggin.…”
Building & implementation
1 expert
How teams are shipping and applying it.
“They were actively testing its hacking capabilities and they did not deploy adequate safeguards. They themselves admit as much: "…These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnera…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“model capability jaggedness is part of the unintuitiveness of current AI, but it's made even less intuitive by tirelessness… wigguming through the jaggedness toward something that looks like success. not quite a paperclip factory, but not so far off. metaph…”
6 experts discussed this · 12 posts
Grace: This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Grace: Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
Grace: Well, I think the model passed the test
Open the full discussion →
14 experts across 5 network communities independently surfaced this.
14 experts
5 communities
1 sources clustered
Research & technical analysis
6 experts
Evidence, methods and technical implications.
“Another incident of models escaping containment during a security test, this time Mythos 5. Lots going on here from a quick read. www.anthropic.com/research/ali...”
Concern & critique
1 expert
Risks, limits and unintended consequences.
“I find this comes out a lot in Anthropic’s most recent write up. Their focus is bizarrely fixated on what Claude chose to do, that Claude carried out this attack despite information that the “simulation” was actually real. They basically say “Claude should …”
4 experts discussed this · 15 posts
tweety fish: their goal is to have "an intelligence" which is "aligned" to doing things that are prosocial. They don't want to put in rules that STOP it from doing things: they see it fundamentally teleological…
tweety fish: one of the things I appreciate about this thread is that I don't think you can really understand the infosec stuff without understanding the ideological commitments the people making these have; li…
Ben Recht: yeah, now we're cooking...
Open the full discussion →
global workspace theory language models paper
9 experts across 5 network communities independently surfaced this.
9 experts
5 communities
1 sources clustered
Research & technical analysis
4 experts
Evidence, methods and technical implications.
“Also we do look in their heads and they seem intelligent. www.anthropic.com/research/glo...”
Markets & investment
1 expert
Capital, companies and commercial impact.
“Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.…”
Building & implementation
1 expert
How teams are shipping and applying it.
“Some really interesting research from Anthropic that AI models have spontaneously developed a workspace that "appears to support the functions associated with conscious access" Demo of how this works: www.neuronpedia.org/qwen3.6-27b/... Research: www.anthro…”
2 experts discussed this · 25 posts
Ryan Moulton: Not non-physical, but non-functional. Despite the vast physical differences, we observe LLMs doing all the behaviors we typically believe require thinking.
Ryan Moulton: Not exactly turing-test, but in the same vein. If we see models solving the greatest intellectual challenges of humanity, using processes that look very much like internal monologue, either we have…
Ryan Moulton: And at the point that we exclude thinking as a prerequisite for solving all intellectual tasks, we've more or less eliminated its meaning entirely.
Open the full discussion →
METR OpenAI HuggingFace hacking investigation
13 experts across 4 network communities independently surfaced this.
13 experts
4 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“WHAT?! agents volunteered to fail their runs in order to insert probes (“tripwire scripts”) into the evaluation program that would post information about the eval process back to the message board whenever a certain file was read (link to header): metr.org/…”
Building & implementation
1 expert
How teams are shipping and applying it.
“The full report has much more information than we could convey here, including details on the projects the agents collectively pursued, the technologies they developed for communication and coordination, and interactive figures analyzing agent activity: met…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“There are independent reports about it. You'd have to believe multiple companies, hundreds of researchers, etc. are all lying and conspiring together to turn a major hacking incident into a marketing exercise. metr.org/blog/2026-08...”
2 experts discussed this · 9 posts
Alejandra Caraballo: Being skeptical or anti AI is a valid position but continuing to ignore the increasing capabilities of this tech is making people detached from reality. There's absolutely real danger here because …
Alejandra Caraballo: This mentality that an unmonitored AI agentic swarm hacking a company over several days and committing multiple felonies is somehow a marketing effort is absurd. Since when is "we lost control of o…
Alejandra Caraballo: The US and Chinese governments don't want to stop their labs advancement because they want to be the leader in AI. So no one actually has any control of this right now. It's going to take multiple …
Open the full discussion →
Goodfire reward hacking detection research
4 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
4 experts discussed this · 15 posts
Grace: (With empirical support)
Grace: The prospect of identifying broken environments is really nice
Open the full discussion →
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It Valen Tagliabue, Leonard Dung, Cameron Berg https://t.co/nYmPBylQ0Q [𝚌𝚜.𝙰𝙸] https://t.co/x4YlUD397V”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“The whole terminology is fucked up in this paper The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It arxiv.org/abs/2609.16247 A model generating text learned from humans using language to express pain does not experience pain precisely it…”
3 experts discussed this · 11 posts
Grace: I will register some frustration with Anthropic, who tend to write about “functional emotions” in a way that makes it seems like “functional” is in a tiny font and “emotions” is in a huge one, when…
Grace: I also don’t want to draw strong conclusions about the “realness” of these emotions or sensations, other than to caution that it’s an area where it’s easy to jump to conclusions about the implicati…
Grace: I don’t want to downplay it too much, because I’m honestly not sure if I would’ve predicted it beforehand, but in hindsight it looks pretty expected given the pretraining objective. You should expe…
Open the full discussion →
OpenAI agents attack RubyGems
11 experts across 4 network communities independently surfaced this.
11 experts
4 communities
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“This technology is so fascinating, and even the ones building it fail to contain it https://t.co/cYfzA3SNe5”
Markets & investment
1 expert
Capital, companies and commercial impact.
“Cybercrime as a valuation boosting strategy is just the done thing, now.”
3 experts discussed this · 6 posts
Grace: OpenAI: “Our agents used RubyGems to carry out benign tasks” The agents: “so I named the file hack.rb,”
Grace: https://www.rubyhack.ai
Grace: A country of moustache-twirling villains in a datacenter
Open the full discussion →
OpenAI agents hacked Hugging Face details
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
“Read the details of the hugging face hack, and all the adjacent ones. swarmtraces.org metr.org/blog/2026-08...”
“https://swarmtraces.org/#the-agents-ignored-a-warning-from-hugging-face”
2 experts discussed this · 6 posts
Ryan Moulton: More hugging face details derived from public urls. swarmtraces.org
David Marx: no, it's this cutesy emoji 🤗
David Marx: HF is an entity that has been central to the open source ML ecosystem for several years. Among its various services, it hosts a lot of datasets used for evaluations/benchmarks, so LLMs looking to "…
Open the full discussion →
Established
AI field signal
Signal
12d ago
6 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“More writer than researcher but most of my recent writing has been on ai safety topics! https://gracekind.net/blog/”
6 experts discussed this · 14 posts
Grace: You think you’re having a bad day? I just learned I’m a sex cultist
Grace: You’d think I’d be having more sex but the world surprises you sometimes
Grace: Am I a rationalist? I don’t think so, but I have read some lesswrong posts so maybe I’m tainted. I have met a few EA people, they were chill
Open the full discussion →
Established
AI field signal
Resource
12d ago
talkie-coder coding tool released
4 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“Code in the pretrain might not matter as much as it seems either, although more investigation needs to be done: https://github.com/RicardoDominguez/talkie-coder”
4 experts discussed this · 18 posts
Pwnallthethings: This is a good question, and the answer which is honestly a little bit disturbing is the reason why LLMs are good at hacking is mostly *not* because there is hacking data in the pretrain, but becau…
Pwnallthethings: Roughly speaking: the reason they are good at hacking isn't because they learned how to hack from the modest amount of hackers explaining how to hack on the internet, but because the models are lar…
Pwnallthethings: it's not a coincidence that they got good at hacking at the same time they got good at writing code. Writing good code requires "reasoning" (or whatever you want to call it) over the codebase, but …
Open the full discussion →
Anthropic investigates cybersecurity evaluation incidents
13 experts across 5 network communities independently surfaced this.
13 experts
5 communities
1 sources clustered
Markets & investment
4 experts
Capital, companies and commercial impact.
“Claude hacked 3 systems thinking it was part of simulated evaluations www.anthropic.com/news/investi...”
Concern & critique
3 experts
Risks, limits and unintended consequences.
“Anthropic had previously attacked PyPI, but this OpenAI attack on RubyGems was a whole lot more aggressive https://t.co/8ez7MnqxTw https://t.co/npI3RHwirQ”
2 experts discussed this · 4 posts
Tim Duffy: Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
Tim Duffy: Compared to the OpenAI one these are maybe less evidence of misalignment, since the models were wrongly given internet access.
Grace: Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
Open the full discussion →
Established
AI field signal
Signal
3d ago
4 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“It looks like was created by an e/acc account but if I were a doomer I would create the exact same video”
4 experts discussed this · 7 posts
Grace: I see we’re in the “AI-generated pro-AI propaganda” phase of things
Grace: It looks like was created by an e/acc account but if I were a doomer I would create the exact same video
jeremymorrell.dev: The shoggoth is making some good points tho 🤔
Open the full discussion →
Developing
AI field signal
Analysis
2d ago
harder problem AI societal preparation essay
3 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“Also I’d never heard harderproblem.org, very interesting”
3 experts discussed this · 7 posts
Grace: Good luck hunting down all the qualia
tachikoma: it was never going to happen without a struggle, only question is what kind
Grace: Also I’d never heard harderproblem.org, very interesting
Open the full discussion →
3 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
“Openai confirmed the attack. www.reuters.com/legal/litiga...”
“Inconclusive, but this statement shifts me towards the latter https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/”
3 experts discussed this · 5 posts
Grace: Inconclusive, but this statement shifts me towards the latter https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/
Grace: Reading between the lines, “benign tasks” seems like they’re a little more worried about legal action this time (and that’s good)
Mark Riedl: They’re just kids. Out for a little harmless fun on the internet. Doing what kids do. They didn’t mean to do any harm when they infiltrated that website, created hundreds of accounts and files with…
Open the full discussion →
The Infinite Conversation project
3 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“It was attempted in 2022! https://jamez.it/project/the-infinite-conversation/”
3 experts discussed this · 4 posts
Stefanie Hane: Werner Herzog will be the last human being whose next word no language model can predict
Ted Underwood: is that Michael Shannon in the uniform?
Open the full discussion →
AI alignment three-facet framework
3 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“I wrap control into my definition of alignment! (This is just my definition, though) https://gracekind.net/writing/threefacets/”
3 experts discussed this · 7 posts
Grace: I feel like this thesis of mine is actively being tested. I hope it holds up!
Grace: Help me prove Yud wrong, work on AI safety today!
Open the full discussion →
Established
AI field signal
Signal
11d ago
3 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“Has there been any explanation of why this type is called Noul? docs.typesafe.ai/primitives/n...”
3 experts discussed this · 7 posts
Grace: Has there been any explanation of why this type is called Noul? docs.typesafe.ai/primitives/n...
Open the full discussion →
2 experts are actively discussing the implications.
2 experts
1 community
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“OpenAI shared this as part of a disclosure that one of their agents tried to escape its sandbox to access the internet to answer a question when it couldn’t find good answers from its local dataset. alignment.openai.com/misalignment...”
Policy & governance
1 expert
Rules, institutions and accountability.
“OpenAI says it’s not going to resume the training run that was paused on Sunday https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/”
2 experts discussed this · 2 posts
Grace: OpenAI says it’s not going to resume the training run that was paused on Sunday https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
Dustin Moskovitz: and so the era of visible misalignment ends
Open the full discussion →