Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,364 searchable experts 2,966 tracked across all sources
Clear

The developments commanding sustained expert attention

One card per development. Sources are clustered; expert reactions remain attributable.

Developing Evaluation & benchmarks Signal 1d ago
⚡ 1 h early

Hugging face model evaluation security incident

18 experts across 6 network communities independently surfaced this.

18 experts 6 communities 1 sources clustered

“Previously, these AI hacking stories were about breaches in test environments, where any question of AI breaching security was purely theoretical. This is something else. openai.com/index/huggin...”

“xkcd.com/2385/ openai.com/index/huggin...”

6 experts discussed this · 11 posts
Grace: This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Grace: Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
Grace: Well, I think the model passed the test
Open the full discussion →
New AI business Research 23h ago
frontier AI business benchmark paper

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Cool paper looking at how AIs solve unbounded, complex business problems in many fields by testing how well they can crack the cases we use to teach MBAs: 1) AI already does extremely well across diverse business topics 2) Models are improving rapidly with …”

New Culture, work & education Research 5h ago
AI automation labor market evaluations paper

Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“MIT research shows AI automation advances as a "rising tide," steadily improving across labor market tasks instead of sudden breakthroughs. By 2030, frontier AI models could achieve 97% success on text tasks, indicating significant workplace changes. https:…”

New Agents & robotics Signal 16h ago

FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Researchers have launched FinanceComplexQA, a new bilingual benchmark enhancing agentic reasoning in financial analysis, with over 2,000 deep research tasks. This tool tackles complex document synthesis, bridging AI capabilities and financial intricacies. h…”

Developing Evaluation & benchmarks Research 1d ago
gaze object prediction benchmark paper

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“A study presents DiSG, a benchmark for OVGOP, enabling models to find unseen targets in real setups. Techniques like text-driven discovery and selective tuning enhance human attention understanding, outperforming prior methods. https://arxiv.org/abs/2607.18827”

Developing Compute & infrastructure Research 1d ago
QuArch LLM computer architecture benchmark

QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“QUARCH is a new benchmark for evaluating LLMs' reasoning in computer architecture, exposing substantial gaps in advanced reasoning despite solid domain knowledge. It aims to enhance AI innovation in system design by deepening LLMs' grasp of architectural pr…”

What experts are discussing without an anchoring article

4 experts · 46 posts · 1d ago
Pwnallthethings: This is going to be a big deal, but easy to get the wrong end of the stick on it. So I think it's worth breaking down what actually happened, why, and what it actually means, based on the public in…
Pwnallthethings: Upfront tl;dr: a model at OAI hacked out of a constrained environment inside OAI and hacked a *different* company autonomously, without authorization from any human in order to creatively solve a t…
Open the thread →