MilaNLP Lab

Why they matter

Researcher with public evidence across NLP & language, AI research.

AI signals
12
past 30d
Sources
4
distinct domains
Discussões
0
past 30d
Latest signal
3d ago
View every signal from MilaNLP Lab →
The Milan Natural Language Processing Group #NLProc #AI milanlproc.github.io

Articles & links

For today's reading group, @marlutz.bsky.social presented "State media control influences large language models" by Waight et al. (2026) Paper: www.nature.com/articles/s41... #NLProc

State media control influences large language models | Nature nature.com
AI Weekly's analysis →
  • Chinese state-media content appears in typical LLM training sets at roughly 41 times the rate of Chinese-language Wikipedia.
  • Across 37 countries, models prompted in the local language produce more regime-favorable responses in countries with lower press freedom.
  • A pretraining experiment with just 6,400 state-scripted documents pushed an open-weight model to pro-government responses nearly 80 percent of the time.
Read full analysis →
View on Bluesky · ♥ 6 ↻ 3 ↩ 0 · 10 from the directory shared this · 94d ago

For today's reading group, @veraneplenbroek.bsky.social presented "Old Habits Die Hard: How Conversational History Geometrically Traps LLMs" by Simhi et al. (2026) Paper: arxiv.org/abs/2603.03308 #NLProc

[2603.03308] Old Habits Die Hard: How Conversational History Geometrically Traps LLMs arxiv.org
AI Weekly's analysis →
  • A new arXiv paper introduces History-Echoes, a framework showing conversation history biases what large language models generate next.
  • Across three model families and six datasets, the authors find gaps in latent space form a 'geometric trap' confining a model's response trajectory.
  • The work suggests hallucinations from earlier turns can influence later responses, implying an early mistake tends to persist inside a session.
Read full analysis →
View on Bluesky · ♥ 11 ↻ 5 ↩ 0 · 2 from the directory shared this · 73d ago

We're back with the reading group! Today Lorena Calvo-Bartolomé presented "The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs" Paper: aclanthology.org/2025.acl-lon... #NLProc #LLMasajudge

aclanthology.org
View on Bluesky · ♥ 7 ↻ 3 ↩ 0 · 2 from the directory shared this · 17d ago

#MemoryModay #NLProc Outstanding Paper at ACL 2024! @paul-rottger.bsky.social et al. evaluate LLM values and opinions in 'Political Compass or Spinning Arrow?' aclanthology.org/2024.acl-lon...

Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models aclanthology.org
AI Weekly's analysis →
  • An ACL 2024 Outstanding Paper argues that LLM political-bias tests using multiple-choice surveys do not reflect how real users query models.
  • In a Political Compass Test case study, models gave substantively different answers when not forced into the test's fixed-choice format.
  • The authors also report that LLM answers shift depending on how models are constrained, and lack paraphrase robustness.
Read full analysis →
View on Bluesky · ♥ 8 ↻ 2 ↩ 0 · 2 from the directory shared this · 20d ago

#MemoryModay #NLProc 'My Answer is C' by Wang et al. (2024) underscores the scrutiny needed for full text responses in LLMs multi-choice evaluations. aclanthology.org/2024.finding...

“My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models aclanthology.org
AI Weekly's analysis →
  • A Findings of ACL 2024 paper reports mismatch rates over 60% between first-token log-probability rankings and the model's actual text answer.
  • The gap holds across final option choice, refusal rate, choice distribution and robustness under prompt perturbation, according to the authors.
  • Models heavily fine-tuned on conversational or safety data are especially impacted, and constraining prompts to force an option letter does not close the gap.
Read full analysis →
View on Bluesky · ♥ 6 ↻ 2 ↩ 0 · 2 from the directory shared this · 62d ago

#MemoryModay #NLProc Fornaciari et al.'s 2022 'Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa' is a useful Python tool for benchmarking NLP predictions with bootstrapped sampling. aclanthology.org/2022.acl-dem...

Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa aclanthology.org
AI Weekly's analysis →
  • BooStSa is an ACL 2022 demo tool that automates bootstrap-sampling significance tests for comparing NLP model results across many conditions.
  • Unlike prior helpers, it handles both standard categorical predictions and soft-label outputs expressed as probability distributions across classes.
  • Tommaso Fornaciari, Alexandra Uma, Massimo Poesio and Dirk Hovy presented the tool at ACL 2022 in Dublin, Ireland.
Read full analysis →
View on Bluesky · ♥ 4 ↻ 3 ↩ 0 · 2 from the directory shared this · 76d ago

For today's reading group, @esradonmez.bsky.social presented "RLHF May Not Reflect Genuine Preferences" by Ghafouri et al. (2026). Interesting thoughts on whether annotations are actually real preferences! Paper: arxiv.org/abs/2604.03238 #NLProc #RLHF

RLHF May Not Reflect Genuine Preferences arxiv.org
AI Weekly's analysis →
  • A new arxiv preprint argues RLHF annotator responses may not represent genuine preferences at all, but responses constructed on the spot.
  • Filtering high-inconsistency annotators in two RLHF datasets flipped majority harm classifications for 18.6% of prompts.
  • The same filtering shifted mean ratings by more than 13 points on a 100-point scale, suggesting systematic rather than random noise.
Read full analysis →
View on Bluesky · ♥ 6 ↻ 1 ↩ 1 · 2 from the directory shared this · 87d ago

#TBT #NLProc Hessenthaler et al.'s 2022 work delves into AI's link with fairness & energy reduction in English NLP models, challenging bias reduction theories. #AI #NLP #sustainability aclanthology.org/2022.emnlp-m...

Bridging Fairness and Environmental Sustainability in Natural Language Processing aclanthology.org
AI Weekly's analysis →
  • An EMNLP 2022 paper reports that knowledge distillation, a common efficiency technique, can actually decrease model fairness rather than preserve it.
  • The case study evaluates distilled models on natural language inference and semantic similarity, with gender bias measured via the Word Embedding Association Test.
  • The authors argue fairness and environmental sustainability are studied in isolation, and that an exclusive focus on one can hinder the other.
Read full analysis →
View on Bluesky · ♥ 7 ↻ 1 ↩ 0 · 2 from the directory shared this · 87d ago

For today's reading group, @pia-p.bsky.social presented "When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning" by Yijiang River Dong et al. (2025) Paper: aclanthology.org/2025.finding... #NLProc

aclanthology.org
View on Bluesky · ♥ 8 ↻ 2 ↩ 0 · 2 from the directory shared this · 115d ago

For today's reading group, @deboranozza.bsky.social presented "LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions" by @myra.bsky.social et al. (2026) Paper: arxiv.org/pdf/2609.14849 #NLProc

arxiv.org
View on Bluesky · ♥ 2 ↻ 1 ↩ 0 · 2 from the directory shared this · 3d ago

#TBT #NLProc 'Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages' funnels limited target-language data to fine-tune models & enhance effectiveness. By @paul-rottger.bsky.social et al. aclanthology.org/2022.emnlp-m...

Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages aclanthology.org
AI Weekly's analysis →
  • Röttger and co-authors find a small amount of target-language fine-tuning data is enough to achieve strong hate speech detection performance.
  • The benefits of adding more target-language annotation decrease exponentially, the EMNLP 2022 paper reports across five non-English languages.
  • Initial fine-tuning on English data can partially substitute for target-language labels and improve model generalisability, per the study.
Read full analysis →
View on Bluesky · ♥ 3 ↻ 1 ↩ 0 · 2 from the directory shared this · 3d ago

Recent commentary

@taniseceron.bsky.social is presenting her work about political content in pre-training and post-training data at the AI & Society conference. #AIandSociety #NLProc

View on Bluesky · ♥ 17 ↻ 5 ↩ 0 · 107d ago

🎉 Excited to welcome Lorena Calvo Bartolomé to our lab! Her expertise in NLP and signal processing will bring fresh perspectives to our research. She will be working on TOLD, a project exploring voice-based data collection as a richer alternative to written annotation. #NLProc

View on Bluesky · ♥ 11 ↻ 5 ↩ 0 · 116d ago

👋 We’re delighted to welcome @celianri.bsky.social, who’ll be visiting our lab over the next few months! Her research explores moderation practices on Reddit, combining social-graph signals with NLP to predict moderation outcomes. Great to have you with us, Célia!

View on Bluesky · ♥ 11 ↻ 3 ↩ 0 · 6d ago

🧠🤖 It was a pleasure to host @andreadevarda.bsky.social for his talk, "Large Language Models as Models of Human Language(s) and Higher-Level Cognition." A truly inspiring talk! #NLProc

View on Bluesky · ♥ 9 ↻ 3 ↩ 0 · 76d ago

We had the pleasure of hosting @tresiwald.bsky.social and Alireza Salemi at our latest seminar. Andreas spoke about the reliability of language models through computational argumentation, while Alireza presented his work on personalizing LLMs. Thank you both for the inspiring talks! #NLProc

View on Bluesky · ♥ 9 ↻ 3 ↩ 0 · 118d ago

We were honored to welcome Giulia Barbareschi to our lab for an inspiring conversation on inclusive technology and what NLP should learn from it. Thank you for sharing your insights and vision! #InclusiveTech #Accessibility #NLProc

View on Bluesky · ♥ 7 ↻ 3 ↩ 0 · 96d ago

Last Friday, we had the pleasure of hosting Clement Jonathan Mazet-Sonilhac and Monika Kaczorowska for an insightful discussion on AI adoption and its impact on the quality of financial services. #NLProc #AI

View on Bluesky · ♥ 6 ↻ 1 ↩ 0 · 109d ago

It was a pleasure to host @radamihalcea.bsky.social at our weekly lab seminar. Thank you for the inspiring talk and thought-provoking discussion on the importance of the long tail in NLP. We really enjoyed the conversation!

View on Bluesky · ♥ 5 ↻ 1 ↩ 0 · 59d ago

In MilaNLP Lab's orbit

Center = MilaNLP Lab. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you MilaNLP Lab? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/milanlp-bsky-social)