Ataraxos tops Stratego champion 15-1-4 for under $8,000
TL;DR
- Ataraxos, from a Carnegie Mellon, MIT, NYU and Stanford team, beat Pim Niemeijer, Stratego's most decorated player, 15-1 with 4 draws across a 20-game series.
- Training cost ran under $8,000 on 16 Nvidia H100 GPUs for one week, roughly 1/500th of DeepMind's DeepNash compute, published in Nature on September 30, 2026.
- The system also beat three Barrage Stratego world champions, set new state-of-the-art marks on Hanabi 2-to-5-player, and outperformed PerfectDou and DouZero on Dou dizhu.
Ataraxos, an AI system from a joint team at Carnegie Mellon, MIT, NYU and Stanford, beat Pim Niemeijer, Stratego's most decorated player, by 15 wins, 1 loss and 4 draws in a 20-game official series. The paper in Nature, published September 30, 2026, puts the training bill at under $8,000.
That budget figure is the real result. The authors report Ataraxos reached higher playing strength than DeepMind's DeepNash while using under 1/100 of the training examples and under 1/30 of the self-play games. "Our system reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency," MIT's Gabriele Farina says in MIT News coverage.
The method pairs self-play reinforcement learning with a belief network that estimates hidden opponent pieces and a test-time search procedure. "With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting," Farina told MIT News.
Ataraxos also defeated three world champions in Barrage Stratego across four 50-game series, achieved new state-of-the-art marks on Hanabi across the 2-to-5-player variants, and outperformed PerfectDou and DouZero on Dou dizhu. Hardware ran to 16 Nvidia H100 GPUs for one week on the reinforcement-learning phase, plus four H100s for four days to train the belief network. Two of the researchers we follow in our Who's Who posted the Nature link within hours of publication.
"Ataraxos is good at calculating risk in a way that humans are not," Farina added. "A human might start freaking out if their most valuable piece is exposed, but the bot can be surprisingly composed."
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature