Ataraxos AI tops Stratego champion for under $8,000
TL;DR
- Ataraxos defeated Stratego's most decorated player, Pim Niemeijer, by 15 wins, 1 loss and 4 draws, with an 85% effective win rate.
- The training run cost under $8,000 on 16 NVIDIA H100 GPUs for a week, roughly 1/500th the compute of prior industrial efforts.
- The same system beat three Barrage Stratego world champions and set new scores in Hanabi and Dou dizhu, which the authors frame as a general design pattern.
The system Ataraxos, from a team at Carnegie Mellon, MIT, Stanford and NYU, defeated Pim Niemeijer, the most decorated player in Stratego's competitive history, by 15 wins, 1 loss and 4 draws, according to a Nature paper published September 30. The authors report an 85% effective win rate against the top human player and say the training run cost under $8,000, roughly 1/500th the compute, 1/30th the games and 1/100th the training examples of prior industrial work on the game.
Stratego has resisted the search-and-self-play recipes that cracked chess and Go because each player cannot see the other's pieces. Ataraxos pairs a policy-value network with a belief network that models hidden information as a generative process, trained on 16 NVIDIA H100 GPUs for one week across 163 million finished games and 208 billion environment steps. "Ataraxos defeated the most decorated human Stratego player of all time by a large margin," the paper reports, framing the result as "the fulfilment of a longstanding aspiration of the field of strategic decision-making."
The same architecture generalises past adversarial play. In the shorter Barrage variant of Stratego it beat three world champions with statistical significance at P < 1.3 × 10⁻⁵. On two-player Hanabi, the cooperative card game that stresses implicit communication, it averaged 24.654 ± 0.007 points with 77.53% perfect games, and it beat state-of-the-art predecessors on Dou dizhu, the three-player Chinese partnership game. Two of the AI researchers in our Who's Who tracker had posted the paper by the day of publication.
"The success of this approach across adversarial, cooperative and team games establishes a design pattern," the authors conclude.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature