Ataraxos AI beats top Stratego champion with 1/500th compute
TL;DR
- Ataraxos beat the most decorated human Stratego player 15 wins to 1 loss with 4 draws, which the authors call the first superhuman Stratego result.
- The system ran on 16 NVIDIA H100 GPUs for a week, about $8,000 of compute, versus the roughly $3-4.5 million spent on prior Stratego AI DeepNash.
- The same recipe produced a superhuman Barrage Stratego AI and state-of-the-art results on cooperative Hanabi and the Chinese card game dou dizhu.
Ataraxos, an AI from a Carnegie Mellon-led team, beat Pim Niemeijer, who the authors call "the most decorated human Stratego player of all time," 15 wins to 1 loss with 4 draws. The result, reported in Nature on 30 September 2026, is "to our knowledge, the first superhuman result in the game's history," and it ran on roughly $8,000 of compute, 16 NVIDIA H100 GPUs for a week.
The point of comparison is DeepNash, the prior Stratego AI, which the authors say cost "$3-4.5 million and 2-3 months on 1,024 TPUv3 nodes." Ataraxos used about 1/500th the compute, 1/30th the games and 1/100th the training examples. Statistical significance against Niemeijer was P < 2.6 × 10⁻⁴.
The system pairs a self-play policy-value network with a belief network that models hidden opponent information, plus a test-time search built on an "update equivalence" framework. "The value of a decision depends not only on the policies agents employ thereafter but also on those they used, and counterfactually would have used, before and at the time of the decision," the authors write. Their recipe uses stronger regularization and larger policy updates early in self-play, then weaker regularization and smaller updates later.
The same techniques produced a superhuman Barrage Stratego AI that swept four 50-game series against three multi-time world champions (aggregate P < 1.3 × 10⁻⁵), plus state-of-the-art results on cooperative Hanabi (24.654 average score in the two-player variant, 24.863 in three-player) and the Chinese card game dou dizhu. The paper frames the method as "a design pattern for reinforcement learning and search that is effective under large amounts of hidden information." Two researchers we track flagged the paper the day it posted.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature