Ataraxos beats Stratego's top pros on under $8,000 compute
TL;DR
- Ataraxos went 15-1-4 against Stratego's most decorated player Pim Niemeijer over 20 games, with training compute under $8,000.
- The same system swept four 50-game series against three two-time Barrage Stratego world champions.
- A single policy-value-belief-search architecture also set new state-of-the-art marks in Hanabi (2-5 player) and Dou Dizhu.
Fifteen wins, one loss, four draws. That was Ataraxos's 20-game record against Pim Niemeijer, described by the researchers as "the most decorated player in history." The compute bill: under $8,000.
The system, trained on 16 NVIDIA H100 GPUs for a week, is described in a Nature paper published September 30 by a team from Carnegie Mellon, MIT, Stanford and NYU. They report an 85% effective win rate against Niemeijer at a P-value under 2.6 × 10⁻⁴, and call it the first superhuman Stratego result by "a large margin." Against three two-time Barrage Stratego world champions across four 50-game series, Ataraxos won every series. At the August 2025 Stratego World Championship it went 38-2 across 40 games against various players.
The sharper number is the comparison to DeepMind's DeepNash, the previous state of the art. The paper reports DeepNash cost "$3,000,000 to $4,500,000" and consumed 5.5 billion self-play games; Ataraxos used 160 million. The authors claim "500x less compute, 30x fewer games, 100x fewer training examples than DeepNash."
The same architecture, a policy-value network paired with a belief network modelling hidden information and a test-time search procedure, produced new state-of-the-art Hanabi scores across the 2 to 5 player variants (24.654 average in two-player, 24.863 in three). In the Chinese card game Dou Dizhu, over 10,000 duplicated deals, it beat the prior best-known system PerfectDou by a role-averaged score of 0.199 ± 0.015.
Human observers of the Stratego matches described Ataraxos's play as "arrogant," with aggressive setups putting high-value pieces near the front. Code and game records are posted at ataraxosai.github.io.
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature