Ataraxos beats top Stratego player on under-$8,000 budget
TL;DR
- Ataraxos beat the most decorated human Stratego player 15-1 with 4 draws over 20 games, trained for under $8,000 on 16 H100 GPUs.
- The system used roughly 1/500th the compute of prior Stratego AI DeepNash, which cost an estimated $3-$4.5 million to train.
- The same recipe produced superhuman play in Barrage Stratego against three world champions, plus state-of-the-art results in Hanabi and dou dizhu.
Ataraxos, a new Stratego AI, beat Pim Niemeijer, described in the Nature paper as 'the most decorated human Stratego player of all time,' by 15 wins to 1 with 4 draws across a 20-game match. Training ran on 16 NVIDIA H100 GPUs for a week and cost under $8,000.
The gap to prior work is the headline the authors lead with. DeepNash, the previous Stratego AI, was trained at an estimated $3 to $4.5 million; Ataraxos used roughly one five-hundredth the compute, one thirtieth the self-play games, and one hundredth the training examples, according to the Nature paper from researchers at Carnegie Mellon, NYU, Stanford and MIT.
The same recipe produced a superhuman AI for Barrage Stratego, beating three multi-time world champions across four 50-game series, and state-of-the-art results on Hanabi and dou dizhu.
The authors frame this as more than a Stratego result. 'The success of this approach across adversarial, cooperative and team games,' they write, 'establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information.' Two researchers we track posted the paper within hours of publication.
The 20-game sample against one human is small, and the paper tests the design pattern only across board and card games. Whether it carries to messier hidden-information domains is a question the paper itself does not pose.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature