WorldCupArena grades 13 AI models on all 104 World Cup matches
TL;DR
- Four systems predicted champion Spain and only the two Claude Opus 4.7 configurations also recovered the exact Spain–Argentina final.
- Claude Opus 4.7 (Thinking + Search) hit 70.7% result accuracy, 1.0 point above a 152-fan panel and 2.4 above BetVictor.
- Adding web search cut overall scores for every tested pair: -0.26 for Claude, -4.26 for GPT-5.4, and -1.03 for Gemini.
A group from Shanghai Jiao Tong, Nanjing, McGill, and University College London published something on Hugging Face this week that is more interesting than the leaderboard framing suggests. WorldCupArena made 13 language models and deep-research agents commit predictions 24 hours before every kickoff of the 2026 FIFA World Cup, then graded them on result and score, likely players and events, match statistics, and the full tournament path across all 104 matches.
The headline result is that the best system barely edges the market. Claude Opus 4.7 (Thinking + Search) reached 70.7% result accuracy, 1.0 percentage point above a 152-person human-fan panel, 2.4 points above BetVictor, and 5.3 points above Polymarket. Its 17.2% exact-score accuracy was 0.9 points above BetVictor. The clearer gain is on the paper's partial-credit Scoreline metric, where Claude beats BetVictor by 15.14 display points because when it misses it tends to miss close.
Four systems predicted champion Spain, both Claude configurations plus GPT-5.4 and Doubao Seed 2.0 Lite, and only the two Claude runs also recovered the exact Spain–Argentina final. All 13 systems named France, Spain, Brazil, and Argentina as semifinalists, so every one of them missed England and wrongly kept Brazil. Perhaps more surprising, adding web search made every tested pair worse on the overall score: -0.26 for Claude, -4.26 for GPT-5.4, and -1.03 for Gemini. The authors are careful to note these are commercial products rather than identical base models with and without search, so they show the observed result but not its cause.
The honest caveat is that the Claude search configuration only completed 58 of the 104 matches while the market baselines cover all of them and the fan panel covers 94, so the head-to-head numbers are descriptive, not paired. The paper also does not answer whether the pattern generalises past football, or why adding search consistently hurt the models it was added to.
What makes the design more useful than a one-off leaderboard is that it is built to be re-run. The same four steps, register a fixture, collect evidence until the deadline, save each model's forecast, then score after the official record lands, can be pointed at a future league or cup, letting newer models be tested without leaning on outcomes already in their training data. That is the piece worth watching.
Originally reported by huggingface.co
Read the original article →Original headline: WorldCupArena Benchmark Tests 13 AI Systems on 104 Matches, Best Model Barely Beats Betting Markets