LAGoM NLP

Why they matter

Researcher with public evidence across AI research, Compute & infrastructure, Culture, work & education.

AI signals
2
past 30d
Sources
1
distinct domains
Discussões
0
past 30d
Latest signal
20d ago
View every signal from LAGoM NLP →
We are the Leuven AI Group of Multilingual NLP (LAGoM NLP), a research lab at the department of Computer Science at KU Leuven, led by @mdlhx

Articles & links

July has been a good month: * Our ACL paper about Wikipedia quality was awarded an SAC highlight (aclanthology.org/2026.acl-lon...) * @colemanhaley.bsky.social joined our lab as a postdoc * Coleman's CoNLL paper about impossible languages won the best paper award (aclanthology…

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP aclanthology.org
AI Weekly's analysis
  • MinHash deduplication removes 28.33% of all non-English Wikipedia articles, mostly from editions known to be dominated by bot-generated content.
  • Cited bot-content shares reach 99% for Cebuano, 90% for Waray and 68% for Swedish Wikipedia, per an Alshahrani et al. 2023 estimate.
  • Language models trained on the filtered Wikipedia largely match or outperform those trained on the raw dumps, with the biggest gains on lower-quality editions.
Read full analysis →
View on Bluesky · ♥ 11 ↻ 1 ↩ 1 · 3 from the directory shared this · 20d ago

July has been a good month: * Our ACL paper about Wikipedia quality was awarded an SAC highlight (aclanthology.org/2026.acl-lon...) * @colemanhaley.bsky.social joined our lab as a postdoc * Coleman's CoNLL paper about impossible languages won the best paper award (aclanthology…

When transformers learn “impossible” languages, what do they learn? aclanthology.org
AI Weekly's analysis
  • Janarthan, Haley and Goldwater train GPT-2 style models on perturbed 'impossible' variants of English and probe them beyond perplexity.
  • On BLiMP minimal pairs the models show only gradual degradation on impossible languages, mediated by information locality.
  • Generation is where the bias lives: the same models produce substantially fewer high-quality sentences at longer lengths.
Read full analysis →
View on Bluesky · ♥ 11 ↻ 1 ↩ 1 · 2 from the directory shared this · 20d ago

In LAGoM NLP's orbit

Center = LAGoM NLP. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.