Ataraxos AI beats Stratego champion 15-1 for under $8,000
TL;DR
- Ataraxos defeated four-time world champion Pim Niemeijer 15-1 with 4 draws over 20 games, an 85% effective win rate.
- Training cost less than $8,000 on 16 H100 GPUs over one week, roughly 1/500th of DeepMind's DeepNash compute spend.
- The same stack (policy-value net, belief net, test-time search) also scaled to Hanabi and the three-player card game Dou Dizhu.
Ataraxos, a system from researchers at Carnegie Mellon, NYU, Stanford and MIT, won 15 of 20 Stratego matches against Pim Niemeijer, the most decorated player in the game's history. The authors report the result in Nature and estimate the training compute at under $8,000.
The gap to the previous state of the art is the striking part. DeepMind's DeepNash, the prior leading Stratego agent, cost somewhere between $3 million and $4.5 million to train by the Ataraxos team's reckoning. The new system comes in at 'roughly 1/500th of the compute cost, 1/30th of the self-play games, and 1/100th of the training examples' versus DeepNash, according to The Decoder. Training ran one week on 16 Nvidia H100 GPUs, plus four more days on four GPUs for a separate belief network.
Niemeijer's record is unusual even by champion standards: four world championships, 15 Dutch national titles, two online world championships, and more than 600 weeks at the top of the world rankings. Against Ataraxos he took one win and four draws across the 20 games, for an 85 percent effective win rate for the machine.
The method is a general recipe for imperfect-information games, built from three pieces: a policy-value network trained through self-play, a belief network that models the hidden state, and a search procedure that refines the policy at test time. The authors also ran it on Hanabi, a fully cooperative card game, and Dou Dizhu, a three-player card game. Two researchers on our Who's Who list circulated the paper on publication.
'The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum,' the paper concludes.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature