Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,364 searchable experts 2,966 tracked across all sources
Clear
Showing signals surfaced by Leshem (Legend) Choshen @EMNLP ×

The developments commanding sustained expert attention

One card per development. Sources are clustered; expert reactions remain attributable.

Developing AI research Research 1d ago
⚡ 26 h early
automated discovery no universal harness paper

Automated Discovery Has No Universally Superior Harness

2 experts are actively discussing the implications.

2 experts 2 communities 1 sources clustered

“Paper : arxiv.org/abs/2607.18235 Github : github.com/akshat57/har...”

“Akshat Gupta, Jermaine Lei, Alexander Lu, Gopala Anumanchipalli, Leshem Choshen Automated Discovery Has No Universally Superior Harness https://arxiv.org/abs/2607.18235”

2 experts discussed this · 15 posts
Leshem (Legend) Choshen @EMNLP: OpenEvolve underperforms simple autmated discovery harnesses the rest are insignificant from each other. The best choice changed across model–problem pairs. We ran a controlled study (3m+ rollouts)…
Leshem (Legend) Choshen @EMNLP: Automated discovery has high run-to-run variance, yet harnesses are often evaluated with only 3–5 runs. When a harness performs better, how do we know it is genuinely better—and not simply lucky?
Leshem (Legend) Choshen @EMNLP: To answer this, we systematically evaluated 30 budget-matched harnesses across 12 model–problem pairs, using repeated-trial statistical analysis. We found :
Open the full discussion →
Established AI research Research 12d ago
⚡ 401 h early
tokenizer language identification paper

What Language is This? Ask Your Tokenizer

2 directory members surfaced this signal.

2 experts 1 community 1 sources clustered

“Effective language identification based on a tokenizer UnigramLM tokenizer already gives probabilities, testing those to identify a language is fast and effective. Whiceh leads me to wonder, can we identify language during training and affect behavior? arxi…”

Developing Models & releases Signal 1d ago

AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms — Google DeepMind

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Searching iteratively for better solutions, for example, you have a code for filling a square with circle (fancy impressive example huh?) then you keep improving your code to make them better. Example: deepmind.google/blog/alphaev...”

Established AI field signal Resource 19d ago
vamsin07 Whisper ASR tools release

GitHub - vamsin07/multilingual-bpe-tokenizers: Whisper-compatible per-lang byte-level BPE tokenizers (recipe + samples) · GitHub

2 experts are actively discussing the implications.

2 experts 1 community 2 sources clustered

“Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...”

2 experts discussed this · 4 posts
Leshem (Legend) Choshen @EMNLP: What they did? Cleaned data, took a big open speech model (whisper) changed the tokenizer and fine-tuned per data.
Leshem (Legend) Choshen @EMNLP: Let's help them: there's a lot more data that needs cleaning, if we do we can train on a scale larger of models and create much much better models (GPU we'll figure out). Ones that are likely to be…
austegard.com: Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...
Open the full discussion →
Established AI field signal Signal 19d ago

babylm.github.io

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“This also raises the differences between the two (hypotheses in pic). Why do models need so much more data for a similar result? Why can humans learn much more efficiently (better is debatable)? (Known as the babyLM challenge c.f. babylm.github.io if you're…”

What experts are discussing without an anchoring article

3 experts · 6 posts · 8d ago
Leshem (Legend) Choshen @EMNLP: I've been told multilingual models are worse in Arabic, but not that they are worse at understanding a simple (English?) table, or understand columns or strategize. It is like all I've been told ab…
Leshem (Legend) Choshen @EMNLP: Is Gemma 4 multilingual? Not really🤖 A true multilingual LLM should share language-invariant skills like spatial understanding across languages. It doesnt😱 We had LLMs play 2D board games against t…
Open the thread →