Ataraxos AI beats world Stratego champion 15-1-4 in Nature
TL;DR
- Ataraxos defeated the most decorated Stratego player of all time 15-1-4 over a 20-game series, the first superhuman result at the top level.
- The system matched DeepMind's prior Stratego AI DeepNash with less than one hundredth of the training examples and one thirtieth of the self-play games.
- The same recipe produced a superhuman AI for Barrage Stratego and state-of-the-art agents for Hanabi and dou dizhu.
Ataraxos, a reinforcement-learning system from researchers at Carnegie Mellon, NYU, Stanford and MIT, defeated the world's most decorated Stratego player 15 wins to 1 with 4 draws over a 20-game series, according to a paper published in Nature on September 30. The system also posted a 39-2 record at the Stratego world championship.
The margin of efficiency is what distinguishes the result from DeepMind's prior Stratego AI, DeepNash. Co-author Gabriele Farina of MIT told MIT News that Ataraxos "reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples and less than one thirtieth of the self-play games." The architecture bundles three components: a policy–value network trained through self-play, a belief network that models hidden information, and a search procedure that refines the policy at test time.
Stratego resists the techniques that cracked poker. "With Stratego, there is an explosion of possible universes you might have to deal with," Farina said.
The same recipe generalized. The paper reports a superhuman AI for Barrage Stratego and state-of-the-art agents for Hanabi and dou dizhu, all with "low cost and high sample efficiency," per the Nature abstract.
Whether that recipe leaves the game tree is a separate question. "Before adoption can happen, we need a way to audit the model's decisions," Farina said.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature