paper web signal

WorldCupArena grades 13 AI models on 104 FIFA 2026 matches

TL;DR

  • WorldCupArena evaluated 13 language models and deep-research agents on all 104 matches of the 2026 FIFA World Cup, benchmarking against bookmakers and human fans.
  • Claude Opus 4.7 with search led at 70.7% result accuracy and 17.2% exact-score accuracy, edging a 68.3% BetVictor odds baseline.
  • Only two of the 13 systems predicted the exact Spain-Argentina final and four correctly picked Spain as champion.

The most interesting thing in the new WorldCupArena paper on arxiv is not the leaderboard, it is what the leaderboard says about how much room a modern AI system actually has to beat a betting market on well-priced questions. The benchmark was posted the day after Spain beat Argentina in the FIFA 2026 final, and it evaluated 13 language models and deep-research agents on all 104 matches of the tournament.

The setup is what makes this worth reading. Before each kickoff, a model either received a standardised evidence pack or was allowed to search the web on its own, then locked in predictions for the result, the exact score, likely players and events, match statistics, and how the competition would ultimately unfold. Predictions were scored against reality once the match ended. The systems on the leaderboard span most of the current frontier: Claude Opus 4.7, GPT-5.4, Gemini 3.1 Pro Preview with and without search, Gemini Deep Research, DeepSeek V4 Pro, Kimi K2.6, MiniMax M2.7, GLM-5.1, Qwen3.7 Max, and Doubao Seed 2.0 Lite.

The numbers are gentle rather than dramatic. Claude Opus 4.7 in the thinking-plus-search configuration led at 70.7% result accuracy and 17.2% exact-score accuracy. The BetVictor odds baseline came in at 68.3% and 16.3%. A crowd of 152 football fans, aggregated, actually beat BetVictor on result accuracy at 69.7%. Polymarket sat further back at 65.4% and 8.7%. Only two of the thirteen systems predicted the exact Spain-Argentina final pairing, and four picked Spain as champion. In other words the best AI forecaster edged the bookmakers by roughly two percentage points on results and about one point on exact scores, and edged crowd-aggregated fans by less than that.

The honest caveat is a small one but load-bearing. A single 104-match tournament is a thin sample for ranking frontier models against each other, and the paper's own framing is that models with similar result accuracy can still differ sharply on players, events, and statistics. What the reporting does not cleanly separate is how much of the top-model advantage came from being allowed to search the live web, where bookmaker prices are one query away.

If there is a forward-looking point, it is probably that the interesting product surface for LLM sports forecasting is not beating Vegas on winners. It is the finer-grained predictions where the benchmark shows real spread between systems and where the bookmakers themselves offer thinner, less efficient markets.