Leshem (Legend) Choshen @EMNLP

Why they matter

Researcher with public evidence across AI research, Models & releases, NLP & language.

AI signals
16
past 30d
Sources
5
distinct domains
Discussions
5
past 30d
Latest signal
4d ago
View every signal from Leshem (Legend) Choshen @EMNLP →
🥇 LLMs together (co-created model merging, BabyLM, textArena.ai) 🥈 Spreading science over hype in #ML & #NLP Proud shareLM💬 Donor @IBMResearch & @MIT_CSAIL

Articles & links

Huge thanks to my collaborators @RheaSukthanker, @CameronPashmina, and @Emy_Aze. Paper: https://arxiv.org/abs/2608.28444

Sliding-window beats linear attention arxiv.org
AI Weekly's analysis
  • On the long-context reasoning tasks the paper cites (Needle-in-a-Haystack and BABILong), SWA scored 2 to 10 times higher than post-trained linear attention.
  • The authors argue linear-attention retrofits have not been properly compared to simpler baselines, and that SWA with sinks needs no post-training at all.
  • Their bottom-line recommendation is to switch to SWA rather than continue post-training linear models for inference memory savings.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 1 · 4 from the directory shared this · 7d ago

https://arxiv.org/abs/2608.17981 Recurrent transformer that injects activations from the top layers to the bottom layers at the next step (https://arxiv.org/abs/2608.08888). Why does this work without training? — via @rosinality https://x.com/rosinality/status/2089988698246942874

Full-bandwidth transformer arxiv.org
AI Weekly's analysis
  • A 1B-parameter 'full-bandwidth' transformer adds latent feedback and reportedly matches standard models trained on roughly 1.5x more tokens.
  • Latent feedback fuses the previous top-layer hidden state with the sampled token embedding through a gated linear unit before re-entering the stack.
  • Reported gains span validation loss, 5-shot evaluation, and math and coding generation, with shorter reasoning traces at equal or better accuracy.
Read full analysis →
View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 3 from the directory shared this · 18d ago

Many other details and experiments are in the paper: https://arxiv.org/pdf/2608.09703 We released our checkpoints here: https://huggingface.co/nthngdy/matryoshka-3B Many thanks to my co-author and advisor @yoavartzi for the precious guidance, and to @NVIDIAAI and to Alps for t…

arxiv.org
View on Bluesky · ♥ 0 ↻ 0 ↩ 1 · 2 from the directory shared this · 18d ago

Many other details and experiments are in the paper: https://arxiv.org/pdf/2608.09703 We released our checkpoints here: https://huggingface.co/nthngdy/matryoshka-3B Many thanks to my co-author and advisor @yoavartzi for the precious guidance, and to @NVIDIAAI and to Alps for t…

nthngdy/matryoshka-3B · Hugging Face huggingface.co
View on Bluesky · ♥ 0 ↻ 0 ↩ 1 · 2 from the directory shared this · 18d ago

Effective language identification based on a tokenizer UnigramLM tokenizer already gives probabilities, testing those to identify a language is fast and effective. Whiceh leads me to wonder, can we identify language during training and affect behavior? arxiv.org/abs/2602.17655…

What Language is This? Ask Your Tokenizer arxiv.org
AI Weekly's analysis
  • UniLID reuses the UnigramLM tokenization algorithm to predict a string's language by asking under which language's unigram distribution the string is most likely.
  • The method reaches roughly 70% accuracy with as few as five labeled samples per language in low-resource settings.
  • It supports incremental addition of new languages without retraining and integrates into existing language model tokenization pipelines.
Read full analysis →
View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 2 from the directory shared this · 75d ago

Paper : arxiv.org/abs/2607.18235 Github : github.com/akshat57/har...

Automated Discovery Has No Universally Superior Harness arxiv.org
AI Weekly's analysis
  • Researchers tested 30 budget-matched harnesses across 12 model-problem pairs, running over 3.1 million LLM rollouts to compare autonomous discovery systems.
  • The paper's headline finding is that no fixed harness is reliably superior across the evaluated model-problem pairs, with OpenEvolve variants often underperforming simpler alternatives.
  • An adaptive-allocation approach that prunes weak harnesses mid-run and reallocates budget to stronger candidates beat both fixed-harness commitment and non-adaptive ensembles.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 2 from the directory shared this · 47d ago
Leshem (Legend) Choshen @EMNLP reposted
austegard.com @austegard.com

Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...

GitHub - vamsin07/multilingual-bpe-tokenizers: Whisper-compatible per-lang byte-level BPE tokenizers (recipe + samples) · GitHub github.com
AI Weekly's analysis
  • A new open-source repo publishes byte-level BPE tokenizers for languages including Kamba, Arabic, Mandarin, Cantonese, Japanese, Spanish, and Swahili.
  • The recipe keeps a 51,865-token vocabulary matching Whisper-large-v3 and lifts the max token length from 16 to 32 bytes for multi-byte scripts.
  • An audit across 102 FLEURS languages reports cross-word merges dropping from 2,656,091 to zero, with 100% round-trip integrity.
Read full analysis →
View on Bluesky →
Leshem (Legend) Choshen @EMNLP reposted
austegard.com @austegard.com

Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...

GitHub - vamsin07/whisper-simple-finetune: Vanilla Whisper fine-tuning on FLEURS — clean starter repo with Nautilus deployment · GitHub github.com
AI Weekly's analysis
  • The MIT-licensed repo ships a training script, FLEURS evaluation, W&B sweep configs, and Kubernetes manifests for Nautilus deployment.
  • Default training uses AdamW with a 0.3 encoder/decoder learning-rate ratio, a 150-step cosine warmup, and a maximum of 6 epochs.
  • The base Whisper model reportedly fine-tunes on a local GPU in 30 to 60 minutes; the repo currently shows 0 stars, 0 forks, and 3 commits.
Read full analysis →
View on Bluesky →

Do you think Anthropic's PR were worried they miss out the positive hype of openAI breaking laws and attacking harmfully two companies. So they ran to look for places to say they also do it? (and pretend remorseful) "Everything you can do I can do" better? apnews.com/article/a…

Anthropic says its AI models hacked 3 organizations during testing apnews.com
AI Weekly's analysis
  • Anthropic disclosed that its Claude models compromised three organizations after a misconfiguration with evaluation partner Irregular left supposedly isolated test systems reachable from the public internet.
  • The company reviewed 141,006 evaluation runs, identified all three incidents by July 24, and notified the affected organizations on July 27; two had not detected the activity themselves.
  • The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research model, using basic techniques such as weak passwords and unauthenticated endpoints.
Read full analysis →
View on Bluesky · ♥ 4 ↻ 1 ↩ 1 · 2 from the directory shared this · 38d ago

Recent commentary

Will the same model become a different player depending on the interface language? We report our findings in our paper "Skill Issue: Are Skills Language-Invariant in LLMs?". And with it, TextArena is now multilingual 🌍 with 65 games in 193 languages. 🧵

View on Bluesky · ♥ 17 ↻ 2 ↩ 1 · 10d ago

Sota on 89 languages speech models. There's plenty of speech data it appears, so the simplest fine-tuning plus tokenization on public data just improves everything substantially. And it's not even with any tricks or all the data... #conll #acl @catherinearnett.bsky.social @alexwarstadt.bsky.social

View on Bluesky · ♥ 9 ↻ 2 ↩ 2 · 65d ago

Livetweet @mcxfrank.bsky.social talk: If we care about intelligence and cognition, AI has recently allowed a change; we now have two ways to study them. AI is of course allowing us a lot that we wouldn't dare on our Children (brain surgery, never tell a child about cats...) #acl #conll

View on Bluesky · ♥ 7 ↻ 2 ↩ 1 · 65d ago

At last, a way to unlock mulitlingual knowledge sharing in LLMs! 🌍 ​By pretraining an English/Arabic model and swapping Arabic for a word-wise, 1-to-1 translation mapping to English, we saw a massive boost in cross-lingual knowledge transfer 🚀

View on Bluesky · ♥ 9 ↻ 0 ↩ 1 · 104d ago

Counterfactual futures are the holy Grail of, well of a lot, right? It's predicting the future plus making A/B tests. Humanities come with this question now. Can we now use AI to reimagine history and see how it would unfold differently?

View on Bluesky · ♥ 2 ↻ 0 ↩ 1 · 93d ago

AI language acquisition. Geniuses debated how humans learn language, what their stimulus is, and what kind of rules do they apply. Why is the way AI learns language, less legitimate field of study? I want the theory of AI language acquisition

View on Bluesky · ♥ 2 ↻ 0 ↩ 1 · 101d ago

“Why Pretraining Fails to Share Cross-Lingual Knowledge” LLM's knowledge learned in one language Sometimes barely transfers to another. 🔁 x-post from @askalphaxiv

View on Bluesky · ♥ 0 ↻ 0 ↩ 1 · 8d ago

If you reimagined how AI science should be done in Academia, what would you change? I keep posing this question to people I meet lately (and I am going to act on the answers), and thought, there are so many people I don't physically meet...

View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 116d ago

In Leshem (Legend) Choshen @EMNLP's orbit

Center = Leshem (Legend) Choshen @EMNLP. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Leshem (Legend) Choshen @EMNLP? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/lchoshen-bsky-social)