Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,365 searchable experts 2,967 tracked across all sources
Clear
Showing signals surfaced by Ramon Astudillo ×

The questions experts are actively pulling apart

One card per development. Sources are clustered; expert reactions remain attributable.

Established Models & releases Development 16d ago
⚡ 83 h early
US government suspends Anthropic models

Statement on the US government directive to suspend access to Fable 5 and Mythos 5

34 experts across 6 network communities independently surfaced this.

34 experts 6 communities 1 sources clustered

“Well, this situation is confusing. www.anthropic.com/news/fable-m...”

“This is wild. www.anthropic.com/news/fable-m...”

5 experts discussed this · 40 posts
SE Gyges: us government has rendered it illegal to give access to fable or mythos to foreign nationals ^_^ www.anthropic.com/news/fable-m...
SE Gyges: this means nobody can have either of them btw. because they haven't KYC'd their customers hard enough.
SE Gyges: dario got what he wanted
Open the full discussion →
Established Evaluation & benchmarks Development 8d ago
⚡ 455 h early
OpenAI limits GPT-5.6 rollout government request

Summary of METR's predeployment evaluation of GPT-5.6 Sol

5 experts across 5 network communities independently surfaced this.

5 experts 5 communities 1 sources clustered

“Some quotes about Sol cheating metr.org/blog/2026-06... 👇”

“this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...”

2 experts discussed this · 6 posts
Tim Kellogg: this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...
Ted Underwood: this is crazy. METR couldn’t measure the task time horizons of GPT-5.6-Sol because it kept hacking the test harness ..with actual exploits metr.org/blog/2026-06...
Tim Kellogg: lol right? so badly wanted a hacker that they got one, and it’s not the kind of employee you want around
Open the full discussion →
Established Evaluation & benchmarks Analysis 6d ago
⚡ 5 h early

Kimi K3, and what we can still learn from the pelican benchmark

4 experts across 2 network communities independently surfaced this.

4 experts 2 communities 1 sources clustered

“My notes on Kimi K3, plus some thoughts on what we can still learn from the pelican benchmark even while it becomes further detached from how good the models are at the things that matter (like agentic tool calling across longer conversations) simonwillison…”

“simonwillison.net/2026/Jul/16/... >The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic’s Claude Sonnet series and making it the most expensive model released by a Chinese…”

2 experts discussed this · 3 posts
Simon Willison: My notes on Kimi K3, plus some thoughts on what we can still learn from the pelican benchmark even while it becomes further detached from how good the models are at the things that matter (like age…
Ramon Astudillo: My notes on Kimi K3, plus some thoughts on what we can still learn from the pelican benchmark even while it becomes further detached from how good the models are at the things that matter (like age…
Ramon Astudillo: death of the pelican test >That connection has been mostly severed now. The GPT-5.6 and Claude Fable 5 pelicans are outclassed by GLM-5.2, and much as I love GLM I don’t think that’s a Fable-class …
Open the full discussion →

What experts are discussing without an anchoring article

2 experts · 7 posts · 10d ago
Ramon Astudillo: This makes me think about the two iron rules from the last 15y of Deep Learning: 1. The bigger model the better 2. The more end to end learned the better Most of the failures in trying to improve t…
Ramon Astudillo: 👆There are currently two (now mostly one) victories against this iron rule, having to do with how these rules fail to work at all "scales" 👇
Open the thread →