Stratego AI hits superhuman level, costs thousands not millions
TL;DR
- A new arXiv paper claims a 'vastly superhuman' Stratego agent trained for 'a few thousand dollars' rather than the millions prior efforts spent.
- The authors credit general-purpose self-play reinforcement learning paired with test-time search adapted for imperfect-information games.
- The abstract frames earlier million-dollar Stratego efforts as having failed to reach top human play, positioning this as a step change in both performance and cost.
An arXiv preprint from Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, J. Zico Kolter and Gabriele Farina claims a Stratego agent at a "vastly superhuman level," trained for "merely a few thousand dollars." The authors frame it as a step change against prior Stratego efforts whose budgets ran into the millions and, by their account, still fell short of top human play.
The recipe, per the abstract, is "general approaches for self-play reinforcement learning and test-time search under imperfect information." The paper situates Stratego as a canonical test of "strategic decision making under massive amounts of hidden information," and treats the earlier industrial-budget attempts as the baseline being overturned.
The abstract reports no win rates, no opponent list, and no evaluation protocol. Everything above the headline claim — Elo, games played, who the humans were — is not in the public summary. Readers wanting to assess the "vastly superhuman" framing will need the full paper and, eventually, outside play against ranked humans or other engines.
Shared on Bluesky by 1 AI expert
Originally reported by arxiv.org
Read the original article →Original headline: Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search