Ataraxos defeats Stratego's top player 15-1-4 in Nature
TL;DR
- An AI called Ataraxos has beaten Stratego's most decorated player 15 wins to 1 across 20 games, the first superhuman result in the game's history.
- The system trained on 16 NVIDIA H100 GPUs for one week, roughly 1/500th the compute of DeepMind's earlier DeepNash effort.
- The same recipe produced superhuman play in Barrage Stratego and state-of-the-art agents in Hanabi and dou dizhu.
Ataraxos, an AI built by researchers at Carnegie Mellon, MIT, Stanford and NYU, has become the first system to play Stratego at a superhuman level. It defeated Pim Niemeijer, the most decorated player in the game's history, with 15 wins, 1 loss and 4 draws across a 20-game match.
The paper, published in Nature on September 30, lands in a game where hidden information has long frustrated AI research. 'Even with multimillion-dollar industrial research efforts, top-human-level play at Stratego … has remained beyond the reach of artificial intelligence,' the authors write.
The cost is the twist. Ataraxos trained on 16 NVIDIA H100 GPUs for one week, consuming what the paper describes as roughly 1/500th the compute, 1/30th the self-play games and 1/100th the training examples of DeepMind's DeepNash. The abstract puts it as 'orders of magnitude less compute and data than previous efforts.'
Niemeijer's record is not close. He holds four world championships, 15 Dutch national titles, two online world titles and more than 600 weeks ranked number one, according to The Decoder. 'The best Stratego player ever,' is how George Franka, who has competed in every world championship since 1997, describes him.
The same recipe generalises. Using the same techniques the team built a superhuman AI for Barrage Stratego and state-of-the-art agents for Hanabi and dou dizhu, spanning adversarial, cooperative and team play. The paper frames the combined result as 'a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.'
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature