Ataraxos beats top Stratego player 15-1 for under $8,000
TL;DR
- An AI called Ataraxos beat Dutch Stratego champion Pim Niemeijer 15 wins, 1 loss, 4 draws over a 20-game series, an 85% effective win rate.
- Training ran one week on 16 Nvidia H100 GPUs plus four days on four more, costing under $8,000 — roughly 1/500th of DeepMind's DeepNash compute.
- The same self-play, belief-network and test-time-search recipe hit state-of-the-art on Hanabi (2-5 players), Dou Dizhu, and Barrage Stratego.
Fifteen wins, one loss, four draws. Over a 20-game series, an AI system called Ataraxos beat Pim Niemeijer, the player George Franka calls "the best Stratego player of all time." That is the headline result in a new Nature paper from a team at Carnegie Mellon, NYU, Stanford and MIT.
"Ataraxos defeated the most decorated human Stratego player of all time by a large margin," the paper reports, calling it "the first superhuman result in the game's history" achieved "while consuming orders of magnitude less compute and data than previous efforts." Reinforcement learning ran one week on 16 Nvidia H100 GPUs; a separate belief network took four more days on four GPUs. The Decoder puts the total bill under $8,000 at 2025 prices, versus $3 million to $4.5 million for DeepMind's DeepNash system, or "roughly 1/500th of the compute cost."
Stratego's difficulty is in the hidden setup: more than 10^33 possible opening configurations before anyone moves. Ataraxos pairs a self-play policy-value network with a belief network that models hidden information from public state, then runs test-time search over sampled game states to refine each decision.
The same recipe produced state-of-the-art results on Hanabi across its 2-5 player variants and on the Chinese card game Dou Dizhu, outperforming PerfectDou and DouZero. In Barrage Stratego, it beat three multi-time world champions.
The 85 percent effective win rate against Niemeijer, counting draws as half, comes from that single twenty-game sitting.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature