Ataraxos beats top human Stratego player, 1/500 the compute
TL;DR
- Ataraxos, built by researchers at CMU, MIT, Stanford and NYU, beat the most decorated human Stratego player 15-1 with 4 draws over 20 games.
- The system used roughly 1/500th DeepNash's compute, training on 16 H100 GPUs for a week at approximately a few thousand dollars.
- The same recipe produced superhuman Barrage Stratego, state-of-the-art Hanabi, and defeated PerfectDou at Dou Dizhu.
Ataraxos, a general-purpose AI for games of imperfect information built by a team from Carnegie Mellon, MIT, Stanford and NYU, beat Pim Niemeijer, described as the most decorated Stratego player in history, with 15 wins, 1 loss and 4 draws across a 20-game series. The result, published in Nature on September 30, was reached at roughly 1/500th the computational cost of prior efforts (DeepNash).
The authors call the 85% effective win rate "without precedent at the highest level of human play". Training ran on 16 NVIDIA H100 GPUs for one week, with a separate belief network trained on 4 GPUs for four days; the paper puts the full training cost at approximately a few thousand dollars.
The same recipe transfers. Ataraxos defeated three top-ranked Barrage Stratego world champions across four 50-game series, the first documented superhuman performance in that variant, and beat PerfectDou at Dou Dizhu by a role-averaged margin of 0.199. On cooperative Hanabi, it posted perfect-game rates of 77.53% in two-player and 89.90% in three-player games.
Three components carry the work: a policy-value network trained via self-play, a belief network that predicts opponent piece types from available information, and a search procedure that samples possible game states at test time. The authors credit a dynamic regularization schedule, stronger early and weaker later, with damping the chaotic learning dynamics that imperfect-information self-play tends to produce.
"The success of the AIs presented herein shows that the presence of large amounts of hidden information is no longer prohibitive," the paper concludes.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature