Ataraxos beats top Stratego pro at ~1/500 DeepMind compute
TL;DR
- Ataraxos beat Pim Niemeijer, Stratego's most decorated human player, 15-1 with 4 draws, for an 85% effective win rate the authors call unprecedented.
- Training cost a few thousand dollars, against the $3M to $4.5M DeepMind spent on DeepNash across 1,024 TPUs over two to three months.
- The same method set new state-of-the-art marks on Hanabi, Barrage Stratego and Dou Dizhu, spanning cooperative, adversarial and team games.
Ataraxos, an AI system from researchers at Carnegie Mellon, MIT, Stanford and NYU, beat Pim Niemeijer, "the most decorated human Stratego player of all time" by the paper's description, 15 games to 1 with 4 draws. The authors report in Nature that this is an "85% effective win rate" that is "without precedent at the highest level of human play."
The budget line is the one that will travel. Training cost "a few thousand dollars," against "between US$3,000,000 and US$4,500,000" for DeepMind's earlier Stratego system DeepNash, which ran on "1,024 tensor processing units for 2-3 months." The team describes the gap as "1/500th of the compute cost, 1/30th of the self-play games and 1/100th of the training examples" of that prior effort.
The same approach also set new state-of-the-art marks on Hanabi, Barrage Stratego and the Chinese card game Dou Dizhu, covering cooperative, adversarial and team settings. "The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information," the authors write.
Two of the AI researchers on our Who's Who list circulated the paper within hours of publication.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature