Vilém Zouhar

Why they matter

Researcher with public evidence across AI research, Evaluation & benchmarks, NLP & language.

AI signals
6
past 30d
Sources
3
distinct domains
Discussões
3
past 30d
Latest signal
1d ago
View every signal from Vilém Zouhar →
PhD @ ETH Zürich | working on (multilingual) evaluation of NLP | on the academic job market | go #vegan | https://vilda.net

Articles & links

Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173

Last Translation Benchmark arxiv.org
AI Weekly's analysis →
  • Last Translation Benchmark ships peer-reviewed hard examples, each paired with handcrafted verification rules that flag concrete failure cases instead of assigning a metric score.
  • LTBv1 covers submissions accepted before September 1st 2026, spans texts, images, audio and videos, and is designed to accept ongoing contributions.
  • The authors argue MT is stuck: standard benchmarks are saturating, automatic metrics are gameable, and human evaluation lacks reproducibility and scalability.
Read full analysis →
View on Bluesky · ♥ 45 ↻ 11 ↩ 3 · 3 from the directory shared this · 26d ago

In the cESA annotation protocol, the annotators see multiple outputs at the same time, which they judge with error spans and absolute scores. This is faster and more objective than doing it one by one. It's implemented in Pearmut and used in WMT26. arxiv.org/abs/2607.26640 WMT

Contrastive ESA: Human Evaluation of Multiple Translations at Once arxiv.org
AI Weekly's analysis →
  • cESA shows annotators several translations of the same source at once, has them mark major and minor error spans, then score each 0-100%.
  • The method was validated on an English-to-Japanese comparison of 12 models, with the authors reporting lower annotation time and noise than pointwise evaluation.
  • Unlike contrastive ranking, cESA yields absolute quality judgments that support non-parametric model rankings without post-hoc statistical corrections.
Read full analysis →
View on Bluesky · ♥ 0 ↻ 0 ↩ 1 · 2 from the directory shared this · 16d ago

Last Translation Benchmark is a live paper+dataset and you can still join last-translation-benchmark.vilda.net Massive thanks to all the >250 dataset contributors and @niyatibafna.bsky.social @mukundc2k.bsky.social @maikezufle.bsky.social @pinzhen.bsky.social

Last Translation Benchmark last-translation-benchmark.vilda.net
View on Bluesky · ♥ 8 ↻ 2 ↩ 0 · 2 from the directory shared this · 26d ago

AI is one such tool. Skeptics say that modern science will devolve into mathematics-as-chess, or mathematics-as-poetry [4], but I'm skeptical of these claims because science was never about generating papers. [4] The Future of Human Mathematics. Jacob Hilton. 2026 www.lesswron…

Jacob_Hilton's Shortform — LessWrong lesswrong.com
View on Bluesky · ♥ 1 ↻ 0 ↩ 1 · 1d ago

Recent commentary

Be the reviewer you want (or your AC wants) to have. (The 1 goes to the AI written paper. No thank you for wasting 2 hours of my life. Pleasure reviewing the rest.)

View on Bluesky · ♥ 3 ↻ 0 ↩ 0 · 92d ago

In Vilém Zouhar's orbit

Center = Vilém Zouhar. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Vilém Zouhar? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/zouharvi-bsky-social)