Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,364 searchable experts 2,966 tracked across all sources
Clear
Showing signals surfaced by arxiv cs.CL ×

The developments commanding sustained expert attention

One card per development. Sources are clustered; expert reactions remain attributable.

New Agents & robotics Research 23h ago
⚡ 261 h early
DeepSearch-World self-distillation search agents paper

DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

2 directory members surfaced this signal.

2 experts 2 communities 1 sources clustered

“Xinyu Geng, Xuanhua He, Sixiang Chen, Yanjing Xiao, Fan Zhang, Shijue Huang, Haitao Mi, Zhenwen Liang, Tianqing Fang, Yi R. Fung DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment https://arxiv.org/abs/2607.07820”

Developing AI research Research 1d ago
⚡ 26 h early
automated discovery no universal harness paper

Automated Discovery Has No Universally Superior Harness

2 experts are actively discussing the implications.

2 experts 2 communities 1 sources clustered

“Paper : arxiv.org/abs/2607.18235 Github : github.com/akshat57/har...”

“Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen Automated Discovery Has No Universally Superior Harness https://arxiv.org/abs/2607.18235”

2 experts discussed this · 15 posts
Leshem (Legend) Choshen @EMNLP: OpenEvolve underperforms simple autmated discovery harnesses the rest are insignificant from each other. The best choice changed across model–problem pairs. We ran a controlled study (3m+ rollouts)…
Leshem (Legend) Choshen @EMNLP: Automated discovery has high run-to-run variance, yet harnesses are often evaluated with only 3–5 runs. When a harness performs better, how do we know it is genuinely better—and not simply lucky?
Leshem (Legend) Choshen @EMNLP: To answer this, we systematically evaluated 30 budget-matched harnesses across 12 model–problem pairs, using repeated-trial statistical analysis. We found :
Open the full discussion →
Developing Models & releases Research 1d ago
⚡ 20 h early
reasoning fine-tuning latent policy states paper

Reasoning Fine-Tuning Induces Persistent Latent Policy States

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Abir Harrasse, Michael Lan, Hunar Batra, Fateme Hashemi Chaleshtori, Chaithanya Bandi Reasoning Fine-Tuning Induces Persistent Latent Policy States https://arxiv.org/abs/2607.18532”

“A study reveals that reasoning fine-tuning transforms models by inducing richer latent-policy structures, enhancing multi-step reasoning through improved dynamics. This boost in performance offers fresh insights into how models handle complex reasoning task…”

Developing Agents & robotics Research 1d ago
⚡ 19 h early
evidence-aware RL long-context repetition paper

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Lizhe Fang, Weizhou Shen, Tianyi Tang, Yisen Wang Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning https://arxiv.org/abs/2607.19345”

“Researchers developed GEAR, a method that reduces repetitive copying in long-context reasoning, improving accuracy by up to +4.6 points across benchmarks, highlighting the importance of focused reasoning in AI tasks. https://arxiv.org/abs/2607.19345”

Developing Agents & robotics Research 1d ago
⚡ 19 h early
LLM essay scoring rubric RL paper

Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Xuefeng Jin, Jiashuo Zhang, Teng Cao, Bin Yang Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards https://arxiv.org/abs/2607.19219”

“RLAES leverages reinforcement learning to enhance automated essay scoring and feedback generation, achieving top QWK scores while ensuring high-quality, rubric-based evaluations. This approach promises to revolutionize educational assessments. https://arxiv…”

Developing Models & releases Research 1d ago
⚡ 18 h early
KV cache LLM inference compression paper

C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Chuheng Du, Junyi Chen, Hanlin Tang, Kan Liu, Tao Lan, Lin Qu, Chaoyue Niu, Shengzhong Liu, Guihai Chen, Fan Wu C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference https://arxiv.org/abs/2607.17715”

“C2KV greatly improves large language model inference efficiency via composable, compressed key-value cache reuse, achieving 17× speedup without sacrificing quality. It addresses KV storage and access costs, supporting scalable long-context applications. htt…”

Established Models & releases Research 6d ago
⚡ 43 h early
LLM robustness prediction flip illusion paper

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Yanzhe Zhang, Sanmi Koyejo, Diyi Yang The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context https://arxiv.org/abs/2607.12963”

“Studies show large language models (LLMs) seem robust to irrelevant context but show per-example instability, creating unpredictable prediction flips. Such instability poses risks in real applications, heightening the need for better reliability evaluations…”

Established Evaluation & benchmarks Research 9d ago
⚡ 27 h early
WILDTRACE long-context reasoning benchmark paper

WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Zixin Chen, Peng Liu, Haobo Li, Rui Sheng, Jianhong Tu, Xiaodong Deng, Fei Huang, Kashun Shum, Dayiheng Liu, Huamin Qu WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning https://arxiv.org/abs/2607.09328”

“WILDTRACE transforms long-document reasoning by integrating natural evidence. It emphasizes reasoning capabilities over retrieval. Featuring 481 tasks from actual texts, it highlights gaps in systems, especially in complex reasoning. https://arxiv.org/abs/2…”

Established Models & releases Research 13d ago
⚡ 3 h early
MAESTRO MoE expert pruning paper

[2607.08601] It Takes a MAESTRO To Prune Bad Experts

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Palaash Goel, Ayush Maheshwari, Tanmoy Chakraborty It Takes a MAESTRO To Prune Bad Experts https://arxiv.org/abs/2607.08601”

“MAESTRO uses Markov chain-based pruning to enhance MoE models, achieving up to 10.61% better performance retention under 50% compression and improved consistency across tasks. This breakthrough tackles the memory bottleneck in large language models. https:/…”

Developing AI research Research 1d ago
SIFT self-improving frozen-gate classifier paper

A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification https://arxiv.org/abs/2607.18358”

Developing Models & releases Research 1d ago
convolution for LLMs architecture paper

Convolution for Large Language Models

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Yuchuan Tian, Yingte Shu, Wei He, Shuo Zhang, Tianchen Zhao, Chao Xu, Xinghao Chen, Yunhe Wang, Hanting Chen, Yu Wang Convolution for Large Language Models https://arxiv.org/abs/2607.18413”

Developing Evaluation & benchmarks Research 1d ago
European MMLU multilingual localization paper

Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Pilar S\'anchez-Gij\'on, Susana Valdez, Sof\'ia Calvo Del Barrio, Florence Bellemont, Anna Kokkinidou, Mihai Cristian Brasoveanu Building a European Multilingual Evaluation Dataset: The MMLU Localisation Project within the EMT Network https://arxiv.org/abs/…”

Developing Models & releases Research 1d ago
LLM vulnerability indicators police logs paper

Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Sam Relins, Daniel Birks Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs https://arxiv.org/abs/2607.18446”

Developing Evaluation & benchmarks Research 1d ago
pathology report generation benchmark paper

PathReportEval: A Systematic Benchmark for Pathology Report Generation

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Suryakant Singh, Sejuti Majumder, Beatrice Knudsen, Joel Saltz, Prateek Prasanna PathReportEval: A Systematic Benchmark for Pathology Report Generation https://arxiv.org/abs/2607.18448”

Developing Models & releases Research 1d ago
RL knowledge graph search LLM paper

Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Jia Ao Sun, Hao Yu, Fengran Mo, Zhan Su, Yuchen Hui, Bang Liu, Jian-Yun Nie Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning https://arxiv.org/abs/2607.18481”