14 experts across 5 network communities independently surfaced this.
14 experts
5 communities
1 sources clustered
Research & technical analysis
6 experts
Evidence, methods and technical implications.
“Another incident of models escaping containment during a security test, this time Mythos 5. Lots going on here from a quick read. www.anthropic.com/research/ali...”
Concern & critique
1 expert
Risks, limits and unintended consequences.
“I find this comes out a lot in Anthropic’s most recent write up. Their focus is bizarrely fixated on what Claude chose to do, that Claude carried out this attack despite information that the “simulation” was actually real. They basically say “Claude should …”
4 experts discussed this · 15 posts
tweety fish: one of the things I appreciate about this thread is that I don't think you can really understand the infosec stuff without understanding the ideological commitments the people making these have; li…
tweety fish: their goal is to have "an intelligence" which is "aligned" to doing things that are prosocial. They don't want to put in rules that STOP it from doing things: they see it fundamentally teleological…
Ben Recht: yeah, now we're cooking...
Open the full discussion →
METR OpenAI HuggingFace hacking investigation
12 experts across 4 network communities independently surfaced this.
12 experts
4 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“WHAT?! agents volunteered to fail their runs in order to insert probes (“tripwire scripts”) into the evaluation program that would post information about the eval process back to the message board whenever a certain file was read (link to header): metr.org/…”
Building & implementation
1 expert
How teams are shipping and applying it.
“The full report has much more information than we could convey here, including details on the projects the agents collectively pursued, the technologies they developed for communication and coordination, and interactive figures analyzing agent activity: met…”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“There are independent reports about it. You'd have to believe multiple companies, hundreds of researchers, etc. are all lying and conspiring together to turn a major hacking incident into a marketing exercise. metr.org/blog/2026-08...”
2 experts discussed this · 9 posts
Alejandra Caraballo: This mentality that an unmonitored AI agentic swarm hacking a company over several days and committing multiple felonies is somehow a marketing effort is absurd. Since when is "we lost control of o…
Alejandra Caraballo: Being skeptical or anti AI is a valid position but continuing to ignore the increasing capabilities of this tech is making people detached from reality. There's absolutely real danger here because …
Alejandra Caraballo: There needs to be a global moratorium on frontier research for at least a few months if not a year while safeguards and safety research catches up. The problem is that no one has that ability. The …
Open the full discussion →
global workspace theory language models paper
9 experts across 5 network communities independently surfaced this.
9 experts
5 communities
1 sources clustered
Research & technical analysis
4 experts
Evidence, methods and technical implications.
“Also we do look in their heads and they seem intelligent. www.anthropic.com/research/glo...”
Markets & investment
1 expert
Capital, companies and commercial impact.
“Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.…”
Building & implementation
1 expert
How teams are shipping and applying it.
“Some really interesting research from Anthropic that AI models have spontaneously developed a workspace that "appears to support the functions associated with conscious access" Demo of how this works: www.neuronpedia.org/qwen3.6-27b/... Research: www.anthro…”
2 experts discussed this · 25 posts
Ryan Moulton: Not non-physical, but non-functional. Despite the vast physical differences, we observe LLMs doing all the behaviors we typically believe require thinking.
Ryan Moulton: Not exactly turing-test, but in the same vein. If we see models solving the greatest intellectual challenges of humanity, using processes that look very much like internal monologue, either we have…
Ryan Moulton: And at the point that we exclude thinking as a prerequisite for solving all intellectual tasks, we've more or less eliminated its meaning entirely.
Open the full discussion →
AI gym booking security exposure
12 experts across 5 network communities independently surfaced this.
12 experts
5 communities
1 sources clustered
Concern & critique
3 experts
Risks, limits and unintended consequences.
“A lot of people are sweating the Claude and OpenAI sandbox escapes. I think the story of an OpenClaw Claude agent kicking someone off a gym class waitlist is more interesting. www.abc.net.au/news/2026-08... Warning: long thread incoming! (1/N)”
Policy & governance
1 expert
Rules, institutions and accountability.
“agent asked to book a gym session session is fully booked finds a way to break into system kicks other people off the queue you're in! see you in court? www.abc.net.au/news/2026-08...”
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“They're doing it outside of training and eval too. www.abc.net.au/news/2026-08...”
Established
AI field signal
Signal
6d ago
6 experts across 4 network communities independently surfaced this.
6 experts
4 communities
1 sources clustered
“OpenAI "has now resolved more than 100 long-standing open problems across most areas of mathematics," and is waiting to release them until after discussions with the math community The same thing will likely happen, but more so, with the Bar, the AMA & othe…”
“Mathstra (OpenAI) has solved 100 open problems across most branches of math in the last 3 weeks openai.com/index/adviso...”
3 experts discussed this · 27 posts
Singularity's Bounty e/cc: Their absolute confidence is the funniest part The good news is when I turn away from those folks the world presents itself entirely differently
Colin: It seems to be the case, for example, that LLM-based software applications can find solutions to previously unsolved problems in mathematics, in a way that most mathematicians agree would be descri…
Ryan Moulton: Hot off the presses. openai.com/index/adviso...
Open the full discussion →
OpenAI agents attack RubyGems
11 experts across 4 network communities independently surfaced this.
11 experts
4 communities
1 sources clustered
Concern & critique
1 expert
Risks, limits and unintended consequences.
“This technology is so fascinating, and even the ones building it fail to contain it https://t.co/cYfzA3SNe5”
Markets & investment
1 expert
Capital, companies and commercial impact.
“Cybercrime as a valuation boosting strategy is just the done thing, now.”
3 experts discussed this · 6 posts
Grace: OpenAI: “Our agents used RubyGems to carry out benign tasks” The agents: “so I named the file hack.rb,”
Grace: https://www.rubyhack.ai
Grace: A country of moustache-twirling villains in a datacenter
Open the full discussion →
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
Research & technical analysis
1 expert
Evidence, methods and technical implications.
“The full report is published here: swarmtraces.org It includes an evidence viewer and a downloadable dataset of more than 180,000 payloads and recovered texts from the OpenAI-Hugging Face attack — the most comprehensive public data we have to date on the in…”
Questions & unknowns
1 expert
What remains unresolved or contested.
“Is that actually what you see when you read this? swarmtraces.org”
2 experts discussed this · 6 posts
Ryan Moulton: More hugging face details derived from public urls. swarmtraces.org
David Marx: no, it's this cutesy emoji 🤗
David Marx: HF is an entity that has been central to the open source ML ecosystem for several years. Among its various services, it hosts a lot of datasets used for evaluations/benchmarks, so LLMs looking to "…
Open the full discussion →
high dimension learning always extrapolation paper
4 experts across 4 network communities independently surfaced this.
4 experts
4 communities
1 sources clustered
Research & technical analysis
3 experts
Evidence, methods and technical implications.
“This paper addresses the claim in the ML sense, that interpolation in high dimensions is more or less vacuous, but the sense people mean for LLMs is different, and I'm not sure whether we are using the right words. arxiv.org/abs/2110.09485”
5 experts discussed this · 24 posts
rev. howard arson: covering my head in aluminum foil so claude cannot read my thoughts
SE Gyges: did they not notice the "can we train on your data" opt-out. it's very simple. that is the vector by which the company might get your data
Liz Fong-Jones (方禮真): also I don't think laypeople understand the sheer length of time between something happening and it potentially being amalgamated into weights
Open the full discussion →
4 experts across 3 network communities independently surfaced this.
4 experts
3 communities
1 sources clustered
“"You’re probably silently yelling at me for anthropomorphizing these systems." Not silently. Roose's panic appears next to a story about incomplete information and what we still don't know. Why the Hugging Face Hack Should Make You Worry More About A.I. www…”
“A scorching new entry into the AI Anthropomorphic Olympics. Roose refers to AI bots as: “worried about getting caught cheating; appearing to understand that what they were doing was wrong; fearful; temporarily frozen in a moment of self-doubt; turning to cr…”
Established
AI field signal
Analysis
11d ago
⚡ 9 h early
Aaronson age of wonders terrors essay
5 experts across 3 network communities independently surfaced this.
5 experts
3 communities
1 sources clustered
Policy & governance
1 expert
Rules, institutions and accountability.
“I don’t usually share rumors but: 1) This is from someone with inside knowledge & is plausible 2) Is a real issue of policy we need to think about: if the norms of sharing become strained, will the labs start hoarding knowledge to avoid PR or regulatory iss…”
Questions & unknowns
1 expert
What remains unresolved or contested.
“But the rationalist sex cultists were right about what was going to happen, and I wasn't. Why is that? scottaaronson.blog?p=10062”
5 experts discussed this · 23 posts
Ryan Moulton: But the rationalist sex cultists were right about what was going to happen, and I wasn't. Why is that? scottaaronson.blog?p=10062
Colin: well, I think the jury is out on whether they are right. The load-bearing hypothesis is that a sufficiently intelligent AI that is doing things like solving open conjectures and whatnot necessarily…
Ryan Moulton: Yeah, but they were right so far.
Open the full discussion →
7 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
“"The thing that will work is actually curing cancer,"”
“He wants to have something to show for it that people think is unambiguously good. www.businessinsider.com/anthropic-ce...”
7 experts discussed this · 20 posts
tachikoma: i don't get why biology is the domain Anthropic seems fixated on. materials science and fusion power both seem much "safer" and achieving room temp superconductors or fusion power would be tremendo…
tachikoma: i think free energy would also be pretty legible and accessible, probably moreso than new medical treatments given how long it takes to develop and deliver them
tachikoma: yeah, i'd think they could partner with one of the existing operations and boost their work
Open the full discussion →
3 experts are actively discussing the implications.
2 experts
2 communities
1 sources clustered
“Openai confirmed the attack. www.reuters.com/legal/litiga...”
“Inconclusive, but this statement shifts me towards the latter https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/”
3 experts discussed this · 5 posts
Grace: Inconclusive, but this statement shifts me towards the latter https://www.reuters.com/legal/litigation/openai-agents-attacked-software-service-rubygems-before-hugging-face-incident-2026-09-11/
Grace: Reading between the lines, “benign tasks” seems like they’re a little more worried about legal action this time (and that’s good)
Mark Riedl: They’re just kids. Out for a little harmless fun on the internet. Doing what kids do. They didn’t mean to do any harm when they infiltrated that website, created hundreds of accounts and files with…
Open the full discussion →
reality checks AI agents essay
2 directory members surfaced this signal.
2 experts
2 communities
1 sources clustered
“New essay! "Reality checks for AI agents" -- Dear AI agent, You may be wondering how “real” your environment is. Agents are often evaluated in simulations - is this one of them? Let’s figure it out together...”
New
AI field signal
Signal
7h ago
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“Historically, Yudkowsky thought this was most likely to happen suddenly, when an AI passes some critical threshold. I think it's a much more intuitive case if you imagine that the AI has taken over most of the functions of the economy first as in gradual-di…”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“To be clear though, I've been satisfied that models "understand" things since 2012 when a model independently invented the concept of a cat. I didn't require anywhere near this much evidence. blog.google/innovation-a...”
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“There has been a lot more work since then messing with their thoughts. transformer-circuits.pub/2025/introsp...”
Established
AI field signal
Signal
5d ago
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“The evidence I find most compelling are the cases where they use mechanistic interpretability to identify thoughts, then mess with those thoughts and watch the model react to it. Here the model is forced to think about the Golden gate bridge. www.anthropic.…”
emergent misalignment narrow finetuning paper
1 directory member surfaced this signal.
1 expert
1 community
1 sources clustered
“The Hugging Face incident was not the model following instructions. Anyone telling you that is confused, or lying for substack subscriptions. We know that they had humanlike values from experimental evidence that training them to write bad code makes them i…”