Leshem (Legend) Choshen @EMNLP

Why they matter

Researcher with public evidence across AI research, Models & releases, NLP & language.

AI signals
9
past 30d
Sources
8
distinct domains
Discusiones
7
past 30d
Latest signal
1d ago
View every signal from Leshem (Legend) Choshen @EMNLP →
🥇 LLMs together (co-created model merging, BabyLM, textArena.ai) 🥈 Spreading science over hype in #ML & #NLP Proud shareLM💬 Donor @IBMResearch & @MIT_CSAIL

Articles & links

Effective language identification based on a tokenizer UnigramLM tokenizer already gives probabilities, testing those to identify a language is fast and effective. Whiceh leads me to wonder, can we identify language during training and affect behavior? arxiv.org/abs/2602.17655…

What Language is This? Ask Your Tokenizer arxiv.org
AI Weekly's analysis
  • UniLID reuses the UnigramLM tokenization algorithm to predict a string's language by asking under which language's unigram distribution the string is most likely.
  • The method reaches roughly 70% accuracy with as few as five labeled samples per language in low-resource settings.
  • It supports incremental addition of new languages without retraining and integrates into existing language model tokenization pipelines.
Read full analysis →
View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 2 from the directory shared this · 34d ago

Paper : arxiv.org/abs/2607.18235 Github : github.com/akshat57/har...

Automated Discovery Has No Universally Superior Harness arxiv.org
AI Weekly's analysis
  • Researchers tested 30 budget-matched harnesses across 12 model-problem pairs, running over 3.1 million LLM rollouts to compare autonomous discovery systems.
  • The paper's headline finding is that no fixed harness is reliably superior across the evaluated model-problem pairs, with OpenEvolve variants often underperforming simpler alternatives.
  • An adaptive-allocation approach that prunes weak harnesses mid-run and reallocates budget to stronger candidates beat both fixed-harness commitment and non-adaptive ensembles.
Read full analysis →
View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 2 from the directory shared this · 6d ago
Leshem (Legend) Choshen @EMNLP reposted
austegard.com @austegard.com

Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...

GitHub - vamsin07/multilingual-bpe-tokenizers: Whisper-compatible per-lang byte-level BPE tokenizers (recipe + samples) · GitHub github.com
AI Weekly's analysis
  • A new open-source repo publishes byte-level BPE tokenizers for languages including Kamba, Arabic, Mandarin, Cantonese, Japanese, Spanish, and Swahili.
  • The recipe keeps a 51,865-token vocabulary matching Whisper-large-v3 and lifts the max token length from 16 to 32 bytes for multi-byte scripts.
  • An audit across 102 FLEURS languages reports cross-word merges dropping from 2,656,091 to zero, with 100% round-trip integrity.
Read full analysis →
View on Bluesky →
Leshem (Legend) Choshen @EMNLP reposted
austegard.com @austegard.com

Website: vamsin07.github.io/buzzasr-docs/ Tokenizers: github.com/vamsin07/mul... SFT models: github.com/vamsin07/whi...

GitHub - vamsin07/whisper-simple-finetune: Vanilla Whisper fine-tuning on FLEURS — clean starter repo with Nautilus deployment · GitHub github.com
AI Weekly's analysis
  • The MIT-licensed repo ships a training script, FLEURS evaluation, W&B sweep configs, and Kubernetes manifests for Nautilus deployment.
  • Default training uses AdamW with a 0.3 encoder/decoder learning-rate ratio, a 150-step cosine warmup, and a maximum of 6 epochs.
  • The base Whisper model reportedly fine-tunes on a local GPU in 30 to 60 minutes; the repo currently shows 0 stars, 0 forks, and 3 commits.
Read full analysis →
View on Bluesky →

Sad to see no African representation in the last translation benchmark, despite LLMs and MT being so bad If you know who would be interested or are interested yourself in contributing(coauthoring), please do and share last-translation-benchmark.vilda.net Viz: lvu5.github.io/tr…

Last Translation Benchmark last-translation-benchmark.vilda.net
View on Bluesky · ♥ 4 ↻ 0 ↩ 0 · 2 from the directory shared this · 1d ago

This also raises the differences between the two (hypotheses in pic). Why do models need so much more data for a similar result? Why can humans learn much more efficiently (better is debatable)? (Known as the babyLM challenge c.f. babylm.github.io if you're interested)

babylm.github.io
View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 25d ago

Recent commentary

Sota on 89 languages speech models. There's plenty of speech data it appears, so the simplest fine-tuning plus tokenization on public data just improves everything substantially. And it's not even with any tricks or all the data... #conll #acl @catherinearnett.bsky.social @alexwarstadt.bsky.social

View on Bluesky · ♥ 9 ↻ 2 ↩ 2 · 24d ago

Livetweet @mcxfrank.bsky.social talk: If we care about intelligence and cognition, AI has recently allowed a change; we now have two ways to study them. AI is of course allowing us a lot that we wouldn't dare on our Children (brain surgery, never tell a child about cats...) #acl #conll

View on Bluesky · ♥ 7 ↻ 2 ↩ 1 · 25d ago

At last, a way to unlock mulitlingual knowledge sharing in LLMs! 🌍 ​By pretraining an English/Arabic model and swapping Arabic for a word-wise, 1-to-1 translation mapping to English, we saw a massive boost in cross-lingual knowledge transfer 🚀

View on Bluesky · ♥ 9 ↻ 0 ↩ 1 · 63d ago

Counterfactual futures are the holy Grail of, well of a lot, right? It's predicting the future plus making A/B tests. Humanities come with this question now. Can we now use AI to reimagine history and see how it would unfold differently?

View on Bluesky · ♥ 2 ↻ 0 ↩ 1 · 52d ago

AI language acquisition. Geniuses debated how humans learn language, what their stimulus is, and what kind of rules do they apply. Why is the way AI learns language, less legitimate field of study? I want the theory of AI language acquisition

View on Bluesky · ♥ 2 ↻ 0 ↩ 1 · 60d ago

If you reimagined how AI science should be done in Academia, what would you change? I keep posing this question to people I meet lately (and I am going to act on the answers), and thought, there are so many people I don't physically meet...

View on Bluesky · ♥ 0 ↻ 0 ↩ 0 · 75d ago

In Leshem (Legend) Choshen @EMNLP's orbit

Center = Leshem (Legend) Choshen @EMNLP. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.