Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,397 searchable experts 4,280 tracked across all sources
Filter the conversation Who is saying what?

Combine a professional role with a reaction lens. Both must match the same attributed contribution.

Clear all
Active evidence filter

Showing developments with attributable Research & technical analysis reactions.

New Network reaction maps

See how experts are reacting—not just what they shared.

Posts are grouped by conversation and tone. Select a lens to filter the stream; these are never permanent labels on people.

Showing signals surfaced by Leshem (Legend) Choshen @EMNLP ×

The developments commanding sustained expert attention

One card per development. Sources are clustered; reaction bundles describe these posts, never the people behind them.

Established AI research Research 7d ago
sliding window beats linear attention

Sliding-window beats linear attention

4 experts across 3 network communities independently surfaced this.

Why this matches Research & technical analysis reaction 3 attributable expert contributions · Leshem (Legend) Choshen @EMNLP, Sung Kim, Alexia Jolicoeur-Martineau
“Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: https://arxiv.org/abs/2608.28444” evidence ↗
4 experts 3 communities 1 sources clustered
How the network is reacting 3 experts are independently emphasizing research & technical analysis.
Reaction converging

Research & technical analysis

3 experts

Evidence, methods and technical implications.

“Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: https://arxiv.org/abs/2608.28444”

2 experts discussed this · 2 posts
Alexia Jolicoeur-Martineau: Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators Rhea Sukt…
Miguel Alonso Jr.: Simple beats complicated: We show that switching to a sliding-window attention mask with attention sinks (at no cost) beats linear attention post-training. Huge thanks to my collaborators Rhea Sukt…
James MacGlashan: Makes me wonder if you can remove the sinks if you use softmax-1?
Open the full discussion →
Established Models & releases Signal 18d ago
⚡ 128 h early

Full-bandwidth transformer

3 experts across 2 network communities independently surfaced this.

Why this matches Research & technical analysis reaction 2 attributable expert contributions · Leshem (Legend) Choshen @EMNLP, Sung Kim
“https://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (https://arxiv.org/abs/2608.08888). Why does this work without training? — via @rosinality https://x.com/rosinality/sta…” evidence ↗
3 experts 2 communities 1 sources clustered
3 experts discussed this · 5 posts
Sung Kim: At decoding time, when you feed **previous hidden state** into the input together with token embedding, and it boosts performance for free. It unlocks the ** full bandwidth ** of the transformer: -…
Sung Kim: - It exposes past information across *depth* to the current layer/position Paper: Full-bandwidth transformer ( arxiv.org/abs/2608.08888 )
SE Gyges: u like can't backprop this tho because your path length increases without limit
Open the full discussion →
Established Models & releases Signal 10d ago
⚡ 31 h early

Skill Issue: Are Skills Language-Invariant in LLMs?

2 directory members surfaced this signal.

Why this matches Research & technical analysis reaction 2 attributable expert contributions · Leshem (Legend) Choshen @EMNLP, arxiv cs.CL
“arXiv: https://arxiv.org/abs/2608.25832 alphaXiv: https://alphaxiv.org/abs/2608.25832 HF Paper: https://huggingface.co/papers/2608.25832 Code: https://github.com/TextArena/TextArena” evidence ↗
2 experts 2 communities 1 sources clustered

“arXiv: https://arxiv.org/abs/2608.25832 alphaXiv: https://alphaxiv.org/abs/2608.25832 HF Paper: https://huggingface.co/papers/2608.25832 Code: https://github.com/TextArena/TextArena”

“Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen Skill Issue: Are Skills Language-Invariant in LLMs? https://arxiv.org/abs/2608.25832”

Established AI field signal Signal 18d ago

arxiv.org

2 directory members surfaced this signal.

Why this matches Research & technical analysis reaction 2 attributable expert contributions · Leshem (Legend) Choshen @EMNLP, Tim Kellogg
“Many other details and experiments are in the paper: https://arxiv.org/pdf/2608.09703 We released our checkpoints here: https://huggingface.co/nthngdy/matryoshka-3B Many thanks to my co-author and advisor @yoavartzi for the precious guidance, and to @NVIDIA…” evidence ↗
2 experts 2 communities 1 sources clustered
Established Models & releases Release 18d ago
⚡ 24 h early
matryoshka 3B model released

nthngdy/matryoshka-3B · Hugging Face

2 directory members surfaced this signal.

Why this matches Research & technical analysis reaction 2 attributable expert contributions · Leshem (Legend) Choshen @EMNLP, Tim Kellogg
“Many other details and experiments are in the paper: https://arxiv.org/pdf/2608.09703 We released our checkpoints here: https://huggingface.co/nthngdy/matryoshka-3B Many thanks to my co-author and advisor @yoavartzi for the precious guidance, and to @NVIDIA…” evidence ↗
2 experts 2 communities 1 sources clustered
How the network is reacting 2 experts are emphasizing research & technical analysis.
Shared emphasis

Research & technical analysis

2 experts

Evidence, methods and technical implications.

“Many other details and experiments are in the paper: https://arxiv.org/pdf/2608.09703 We released our checkpoints here: https://huggingface.co/nthngdy/matryoshka-3B Many thanks to my co-author and advisor @yoavartzi for the precious guidance, and to @NVIDIA…”

Established AI field signal Signal 18d ago
⚡ 33 h early

Recirculation

2 directory members surfaced this signal.

Why this matches Research & technical analysis reaction 1 attributable expert contribution · Leshem (Legend) Choshen @EMNLP
“https://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (https://arxiv.org/abs/2608.08888). Why does this work without training? — via @rosinality https://x.com/rosinality/sta…” evidence ↗
2 experts 2 communities 1 sources clustered
Established Evaluation & benchmarks Resource 4d ago
Last Translation Benchmark tool

Last Translation Benchmark

3 experts are actively discussing the implications.

Why this matches Research & technical analysis reaction 2 attributable expert contributions · Leshem (Legend) Choshen @EMNLP, Vilém Zouhar
“Sad to see no African representation in the last translation benchmark, despite LLMs and MT being so bad If you know who would be interested or are interested yourself in contributing(coauthoring), please do and share last-translation-benchmark.vilda.net Vi…” evidence ↗
2 experts 2 communities 1 sources clustered
How the network is reacting 2 experts are emphasizing research & technical analysis.
Shared emphasis

Research & technical analysis

2 experts

Evidence, methods and technical implications.

“Last Translation Benchmark is a live paper+dataset and you can still join last-translation-benchmark.vilda.net Massive thanks to all the >250 dataset contributors and @niyatibafna.bsky.social @mukundc2k.bsky.social @maikezufle.bsky.social @pinzhen.bsky.social”

3 experts discussed this · 9 posts
Vilém Zouhar: There are many things machine translation still can't do. Help us steer the next direction by contributing hard-to-translate inputs (and be on a cool paper).
Vilém Zouhar: Multiple things made us start this effort. Typical translation benchmarks are.. ...oftentimes trivial or saturated (so they can't be used for guiding the next steps in the field) ...not evaluatable…
Vilém Zouhar: In the Last Translation Benchmark we solve both by: - collecting hard-to-translate inputs (texts, images, audios) - requiring human-readable "verification rules", which enable provable evaluation o…
Open the full discussion →
Established Models & releases Research 4d ago
cross-lingual LLM knowledge sharing paper

Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs

1 directory member surfaced this signal.

Why this matches Research & technical analysis reaction 1 attributable expert contribution · Leshem (Legend) Choshen @EMNLP
“arxiv.org/abs/2408.10646 arxiv.org/abs/2502.21228 And so much more (this nuance is somehow not being recognized and following the literature is so hard this days this well established branch on lack of sharing is hardly discussed outside the branch)” evidence ↗
1 expert 1 community 1 sources clustered

“arxiv.org/abs/2408.10646 arxiv.org/abs/2502.21228 And so much more (this nuance is somehow not being recognized and following the literature is so hard this days this well established branch on lack of sharing is hardly discussed outside the branch)”

Established Evaluation & benchmarks Research 4d ago
ECLeKTic cross-lingual benchmark release

ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer

1 directory member surfaced this signal.

Why this matches Research & technical analysis reaction 1 attributable expert contribution · Leshem (Legend) Choshen @EMNLP
“arxiv.org/abs/2408.10646 arxiv.org/abs/2502.21228 And so much more (this nuance is somehow not being recognized and following the literature is so hard this days this well established branch on lack of sharing is hardly discussed outside the branch)” evidence ↗
1 expert 1 community 1 sources clustered

“arxiv.org/abs/2408.10646 arxiv.org/abs/2502.21228 And so much more (this nuance is somehow not being recognized and following the literature is so hard this days this well established branch on lack of sharing is hardly discussed outside the branch)”

Established Models & releases Signal 10d ago

Paper page - Skill Issue: Are Skills Language-Invariant in LLMs?

1 directory member surfaced this signal.

Why this matches Research & technical analysis reaction 1 attributable expert contribution · Leshem (Legend) Choshen @EMNLP
“arXiv: https://arxiv.org/abs/2608.25832 alphaXiv: https://alphaxiv.org/abs/2608.25832 HF Paper: https://huggingface.co/papers/2608.25832 Code: https://github.com/TextArena/TextArena” evidence ↗
1 expert 1 community 1 sources clustered

“arXiv: https://arxiv.org/abs/2608.25832 alphaXiv: https://alphaxiv.org/abs/2608.25832 HF Paper: https://huggingface.co/papers/2608.25832 Code: https://github.com/TextArena/TextArena”

Established Models & releases Signal 18d ago

MatFormer: Nested Transformer for Elastic Inference

1 directory member surfaced this signal.

Why this matches Research & technical analysis reaction 1 attributable expert contribution · Leshem (Legend) Choshen @EMNLP
“We compare to MatFormer (https://arxiv.org/abs/2310.07707), and show that Matryoshka suites offer a better size-performance tradeoff, allow more size flexibility, and deliver variable KV cache requirements In short we trade their inference-time flexibility …” evidence ↗
1 expert 1 community 1 sources clustered

“We compare to MatFormer (https://arxiv.org/abs/2310.07707), and show that Matryoshka suites offer a better size-performance tradeoff, allow more size flexibility, and deliver variable KV cache requirements In short we trade their inference-time flexibility …”

Established Culture, work & education Signal 7d ago

Alexia Jolicoeur-Martineau (@jm_alexia) on X

1 directory member surfaced this signal.

Why this matches Research & technical analysis reaction 1 attributable expert contribution · Leshem (Legend) Choshen @EMNLP
“Linear attention has limited memory, so it must choose which tokens to remember and which to forget. Trying to learn this during post-training at a small cost is misguided. — x-post from @jm_alexia https://x.com/jm_alexia/status/2094414735408050687” evidence ↗
1 expert 1 community 1 sources clustered

“Linear attention has limited memory, so it must choose which tokens to remember and which to forget. Trying to learn this during post-training at a small cost is misguided. — x-post from @jm_alexia https://x.com/jm_alexia/status/2094414735408050687”