Ataraxos beats Stratego champion at 1/500th DeepNash compute
TL;DR
- Ataraxos defeated Stratego's most decorated player Pim Niemeijer 15-1-4 in a 20-game official series, an 85% effective win rate.
- The system trained on 16 NVIDIA H100 GPUs for one week at a cost of a few thousand dollars, roughly 1/500th of DeepMind's DeepNash.
- The same method set new state-of-the-art on Barrage Stratego, 2-to-5-player Hanabi, and Dou dizhu, not just Stratego.
Ataraxos, a reinforcement-learning system from a six-author team at Carnegie Mellon, MIT, NYU and Stanford, defeated the most decorated Stratego player in history, Pim Niemeijer, by 15 wins to 1 loss with 4 draws across a 20-game official series. The authors describe the resulting 85% effective win rate as "unprecedented at the highest level," and they trained the system on 16 NVIDIA H100 GPUs for one week at a cost of a few thousand dollars. The paper runs in Nature's 30 September issue.
The compute delta against DeepMind's DeepNash, the previous Stratego benchmark, is the number carrying the paper: "roughly 1/500th of the compute cost, 1/30th of the self-play games, and 1/100th of the training examples." The method pairs a policy-value network with a second network that infers the likely identity of hidden opponent pieces, plus a test-time search procedure the authors call update equivalence. Lead author Samuel Sokota frames the core difficulty of imperfect-information play directly: "the value of a move depends not only on how play continues afterward but also on what happened before and during the decision."
The same recipe lifted state-of-the-art on games well beyond Stratego, with a first superhuman result on Barrage Stratego, new highs on 2-to-5-player Hanabi, and gains over prior work on the Chinese card game Dou dizhu. A direct head-to-head rematch against DeepNash was not possible; the researchers report that DeepMind indicated "the DeepNash code no longer works."
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature