Ataraxos AI beats top Stratego player at 1/500th the cost
TL;DR
- Ataraxos beat Pim Niemeijer, the most decorated Stratego player in the game's history, 15 wins to 1 loss with 4 draws over 20 games.
- Training ran roughly $8,000 on 16 NVIDIA H100s for a week, about 1/500th the compute cost of DeepMind's prior DeepNash system.
- The same recipe set new state-of-the-art marks on Hanabi for 2 to 5 players, Barrage Stratego, and Dou Dizhu.
Ataraxos, a reinforcement-learning agent built by researchers at Carnegie Mellon, MIT, Stanford and NYU, beat Pim Niemeijer, the most decorated Stratego player in the game's history, 15 wins to 1 loss with 4 draws over a 20-game series. The result, published in Nature on September 30, is described by the authors as the first superhuman result in Stratego, a two-player game where each side keeps its pieces hidden from the opponent.
The second headline is the price. Training Ataraxos ran roughly $8,000 on 16 NVIDIA H100s for one week, plus four more H100s for four days to train its belief network. DeepMind's prior Stratego system, DeepNash, cost between $3 million and $4.5 million. The paper reports "1/500th the compute cost and 1/30th the self-play games."
The same recipe, a policy-value network paired with a belief network modelling hidden information and a test-time search procedure, set new state-of-the-art marks across Hanabi for 2 to 5 players, Barrage Stratego, where it defeated three world champions in 50-game series, and Dou Dizhu, where it beat the previous leader PerfectDou. In a demonstration round at a World Championship, Ataraxos went 38 wins to 2 losses with no draws.
The core trick is a schedule the authors call dynamic regularization coupling: "stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller policy updates late in self-play." That shift, they argue, is what makes self-play tractable when much of the state is hidden from the learner.
Ataraxos trained on 163 million finished games and 208 billion environment steps, running on a custom GPU-accelerated simulator the authors clock at around 10 million board state updates per second. Two researchers in our tracking network posted the paper on publication day.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature