Ataraxos tops Stratego champion at 1/500th DeepNash compute
TL;DR
- Ataraxos defeated Pim Niemeijer, the most decorated human Stratego player, by 15 wins, 1 loss and 4 draws in a 20-game series.
- Training consumed a few thousand dollars and roughly 1/500th the compute of previous industrial efforts like DeepNash.
- The same self-play RL plus test-time search set state-of-the-art on Barrage Stratego, Hanabi 2-5 players, and Dou Dizhu.
Ataraxos, an AI built at Carnegie Mellon, MIT, Stanford and NYU, beat the most decorated human Stratego player of all time, Pim Niemeijer, by 15 wins, 1 loss and 4 draws across a 20-game series. The result is reported in a Nature paper by Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter and Gabriele Farina.
Stratego is a two-player game in which most of the state is hidden. "With Stratego, there is an explosion of possible universes you might have to deal with," Farina told MIT News. The paper calls Ataraxos's 85% effective win rate against Niemeijer "without precedent at the highest level of play."
The compute figure is the other number worth staring at. Training consumed "a few thousand dollars" and about 1/500th the compute cost of previous industrial efforts. "Our system reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples," Farina said. The reinforcement-learning run used 16 NVIDIA H100 GPUs for one week; the belief model used 4 H100 GPUs for four days.
The techniques are not Stratego-specific. The same approach, tabula rasa self-play reinforcement learning plus test-time search with a separate belief network that models hidden information, also beat three multi-time world champions at Barrage Stratego, set state-of-the-art on Hanabi across 2-5 player variants, and beat prior best bots at Dou Dizhu.
Farina is cautious about what comes next. "Before adoption can happen, we need a way to audit the model's decisions. We still have a long way to go." Two researchers on our tracker shared the paper within a day of publication.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature