@UvA_Amsterdam @TheyCallMeMr_ @johnamqdang @mgalle @mziizm @ahmetustun89 @beyzaermis @ananya__sahu @mhrnz_m @mrdanieldsouza @rsrchvslstrytlr @ConorFinlay3 @joshuakurien_ Announcing new research on the state of the art of linguistic reasoning. 🧩 Led by @DanMirea4, Sara Rajaee, …
- Claude Opus 4.8 earned a jury score equivalent to a gold medal on the IOL-AI Challenge, run on unseen 2026 International Linguistics Olympiad problems.
- 14B submissions outperformed models twice their size, with the authors attributing the wins to decoding and output-handling rather than model capacity.
- Automatic metrics ranked systems in the same order as the jury but upscored weak systems by roughly 13 points and understated strong ones.