Ataraxos beats top Stratego player, trained for under $8,000
TL;DR
- Ataraxos defeated Stratego's most decorated player Pim Niemeijer with 15 wins, 1 loss and 4 draws in a 20-game series.
- The system trained for under $8,000 versus $3-4.5 million for prior efforts, using 1/500th the compute of DeepMind's DeepNash.
- The same method set state-of-the-art marks in Barrage Stratego, Hanabi (24.654-24.863/25), and Dou Dizhu over PerfectDou and DouZero.
Ataraxos, a reinforcement-learning system from a team at Carnegie Mellon, NYU, Stanford and MIT, beat Stratego's most decorated player, Pim Niemeijer, with "15 wins, 1 loss and 4 draws" in a 20-game series, and hit an 85% effective win rate against the world's top-ranked player. The result, published September 30, 2026 in Nature, is the paper's headline evidence for superhuman play in a game of imperfect information.
The economics are the second finding. The authors report training the system for under $8,000 and say it consumed "1/500th the compute, 1/30th the games, and 1/100th the training examples compared to DeepMind's DeepNash," the earlier Stratego system whose previous efforts the paper puts at $3-4.5 million.
Ataraxos is not Stratego-specific. The paper reports "First superhuman performance, winning all four 50-game series against top-ranked players" in Barrage Stratego, new state-of-the-art Hanabi scores of 24.654 to 24.863 out of 25 across 2-to-5-player variants, and victories over PerfectDou and DouZero in Dou Dizhu with statistical significance. A demo at the 2025 World Championship produced a 95% effective win rate across 40 games.
The method stitches together three parts: a policy-value network trained by self-play, a separate belief network for the hidden information, and a test-time search that samples hidden states and runs depth-limited rollouts. A choice the authors call "dynamically damped self-play" coordinates regularization strength with policy update size, with "stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller policy updates late in self-play." Training ran on 16 NVIDIA H100 GPUs for one week, plus four H100s for four days on the belief network, consuming 163 million self-play games and 208 billion environment steps.
Two of the researchers we follow flagged the paper on the day it published. The authors report a P-value below 2.6 × 10⁻⁴ for the Niemeijer result under game-independence assumptions.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature