Ataraxos AI beats Stratego champion for under $8,000
TL;DR
- Ataraxos defeated Pim Niemeijer, the most decorated human Stratego player of all time, 15 games to 1 with 4 draws over a 20-game series.
- Training cost under $8,000 at 2025 pricing versus roughly $3-$4.5 million for DeepNash: 1/500th the compute, 1/30th the self-play games.
- Same recipe hit a new state-of-the-art in Hanabi across 2-5 players and beat PerfectDou and DouZero in dou dizhu.
Ataraxos, an AI system built by researchers at Carnegie Mellon, MIT, NYU and Stanford, defeated Pim Niemeijer, described by the paper as "the most decorated human Stratego player of all time," 15 games to 1 with four draws over a 20-game series. The system reached superhuman play in Stratego using what the authors, writing in Nature, call "orders of magnitude less compute and data than previous efforts."
The cost comparison is the headline. The earlier DeepNash agent cost roughly $3 million to $4.5 million to train. Ataraxos cost under $8,000 at 2025 pricing, consuming 1/500th the compute, 1/30th the self-play games and 1/100th the training examples. A separate 40-game World Championship exhibition produced a 95% effective win rate.
The same recipe transferred. In Barrage Stratego, Ataraxos defeated three multi-time world champions. In Hanabi, the authors report a new state-of-the-art across the 2-to-5-player variants. In dou dizhu, it beat the specialist systems PerfectDou and DouZero.
The system has three parts: a policy-value network trained through self-play, a belief network that models hidden information from what each player can see, and a test-time search that refines moves during play. The authors flag one specific lever: coordinating regularization strength with policy update size, applying "stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller policy updates late in self-play." A GPU-accelerated Stratego simulator running at "roughly 10 million board state updates per second" made the compute budget work.
The paper's framing is unusually direct for a Nature result: "reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information."
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature