Ataraxos beats top Stratego player at fraction of DeepNash cost
TL;DR
- Ataraxos defeated Pim Niemeijer, Stratego's most decorated player, with 15 wins, 1 loss and 4 draws across a 20-game series.
- The system was trained for 'a few thousand dollars' on 16 H100 GPUs for one week, versus $3-4.5 million for DeepMind's DeepNash.
- It also beat top Barrage Stratego players, PerfectDou and DouZero at Dou Dizhu, and set new Hanabi scores across 2-5 players.
Ataraxos, an AI system built by researchers at Carnegie Mellon, MIT, NYU and Stanford, defeated Pim Niemeijer — Stratego's most decorated player — with "15 wins, 1 loss and 4 draws" across a 20-game series. The paper, published in Nature on 30 September 2026, calls the result "the first superhuman result in the game's history."
What stands out is the resource gap. The team trained Ataraxos on 16 H100 GPUs over one week, running roughly 160 million self-play games at a cost the authors describe as "a few thousand dollars." DeepMind's prior Stratego system, DeepNash, used around 5.5 billion games at a reported cost of between $3 million and $4.5 million.
Stratego was not the only test. Ataraxos beat three of the four top-ranked Barrage Stratego players — including a "two-time Barrage World Champion" — across four 50-game series, defeated PerfectDou and DouZero at the Chinese card game Dou Dizhu with statistical significance, and set new marks on the cooperative card game Hanabi, scoring 24.654 in two-player play with 77.53% perfect games and 24.863 in three-player with 89.90% perfect games.
The authors — Samuel Sokota and J. Zico Kolter (CMU), Eugene Vinitsky (NYU), Hengyuan Hu (Stanford), Zhiyuan Fan and Gabriele Farina (MIT) — attribute the efficiency to a technique they call "dynamic damping," which coordinates regularisation strength with policy update size during training. Two researchers we follow shared the paper the week it ran.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature