Ataraxos beats Stratego champion on $8,000 training budget
TL;DR
- Ataraxos beat Pim Niemeijer, described as the most decorated Stratego player ever, 15 wins, 1 loss, 4 draws across a 20-game series.
- Training ran on 16 H100 GPUs for one week at roughly $8,000, about 1/500th the compute of DeepMind's DeepNash.
- The same system set state of the art on Hanabi across 2-5 player variants and beat PerfectDou and DouZero at Dou Dizhu.
A team from Carnegie Mellon, NYU, Stanford, and MIT reports in Nature that their system, Ataraxos, beat Pim Niemeijer, described in the paper as the most decorated Stratego player ever, by "15 wins, 1 loss, 4 draws" across a 20-game series.
Training cost: approximately $8,000, using 16 H100 GPUs for one week. That is roughly 1/500th the compute cost of DeepMind's prior Stratego system, DeepNash.
Ataraxos has three parts: a policy-value network trained through self-play, a belief network that predicts opponent piece types from observable information, and test-time search that uses the belief network to sample possible game states before running depth-limited rollouts.
The authors describe the result as establishing "a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field."
The same system also reported first superhuman results at Barrage Stratego, defeating three multi-time world champions across four 50-game series, new state of the art on Hanabi across the 2-to-5-player variants with two orders of magnitude less compute than prior work on the 2-player variant, and wins over PerfectDou and DouZero at Dou Dizhu.
The paper publishes a single headline cost and does not break down which share of the $8,000 was Stratego training versus the other three games, or how many runs were discarded getting there.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature