Ataraxos AI beats Stratego champion Niemeijer 15-1-4
TL;DR
- Ataraxos defeated top-ranked Stratego player Pim Niemeijer 15 wins to 1 loss with 4 draws across a 20-game series.
- Training ran on 16 H100 GPUs for about a week at roughly $8,000, versus $3-4.5 million for the prior DeepNash effort.
- The same system reached superhuman play on Barrage Stratego, Hanabi, and dou dizhu using identical techniques.
An AI system called Ataraxos beat Pim Niemeijer, the most decorated Stratego player in history, 15 wins to 1 loss with 4 draws across a 20-game series, a result described in the paper as "without precedent at the highest level of human play."
The striking number isn't the win column. Training cost about $8,000 on 16 H100 GPUs for a week, against the $3 million to $4.5 million industrial effort behind the earlier Stratego system DeepNash, which never reached superhuman play. The team from Carnegie Mellon, NYU, Stanford, and MIT, writing in Nature, reports the same techniques also beat three multi-time Barrage Stratego champions across four 50-game series, hit state-of-the-art on 2-to-5-player Hanabi, and defeated the previous dou dizhu bots PerfectDou and DouZero with statistical significance.
Three components do the work: a policy-value network, a belief network that models hidden information, and a test-time search that samples possible worlds and evaluates candidate moves. The glue is what the paper calls dynamic damping, with "stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller policy updates late." Human observers described Ataraxos's play as "preternaturally lucky" in piece placement and strong at "long-term positional play."
Two researchers we track shared the paper the day it ran. The 20-game sample against one human is small, and the test-time search itself is modest at 40-ply depth with 1,000 rollouts averaging 1.26 seconds per move. The authors state the broader claim plainly: "reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information."
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature