Ataraxos AI tops Stratego champion for a few thousand dollars
TL;DR
- Ataraxos beat Pim Niemeijer, the most decorated Stratego player of all time, 15-1-4 across a 20-game series.
- The system trained on 16 NVIDIA H100 GPUs for one week, versus an estimated $3-4.5 million the paper puts on DeepMind's DeepNash.
- The same method delivered first superhuman play on Barrage Stratego and improved results on five-player Hanabi and Dou Dizhu.
An AI called Ataraxos beat Stratego's most decorated player, Pim Niemeijer, with 15 wins, 1 loss and 4 draws across a 20-game series, trained for roughly a few thousand dollars of GPU time. The system, published in Nature on 30 September 2026 by researchers from Carnegie Mellon, MIT, NYU and Stanford, trained on 16 NVIDIA H100 GPUs for one week, with the belief network taking a further four H100s for four days.
The authors frame the contribution less as a Stratego scalp than as a method: "general techniques...for both self-play reinforcement learning and test-time search under hidden information." The same recipe produced the first superhuman performance on Barrage Stratego, scored 24.410 ± 0.009 with 58.05% ± 0.49% perfect games on five-player Hanabi, and outperformed the PerfectDou baseline on Dou Dizhu.
The cost comparison is the headline for anyone outside the top labs. The paper estimates DeepMind's earlier Stratego system, DeepNash, would cost $3-4.5 million to run today, and reports that Ataraxos "reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples." Two researchers we follow at AI Weekly circulated the paper the week it dropped.
The authors describe the work as "a design pattern for reinforcement learning and search that is effective under large amounts of hidden information," which they call "a longstanding desideratum of the field." Twenty games against one champion is still twenty games; the broader claim rides on the three other game benchmarks and on whether the method travels to problems with less tidy hidden state.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature