nature.com web signal

Ataraxos beats top Stratego player on under-$8,000 budget

TL;DR

  • Ataraxos beat the most decorated human Stratego player 15-1 with 4 draws over 20 games, trained for under $8,000 on 16 H100 GPUs.
  • The system used roughly 1/500th the compute of prior Stratego AI DeepNash, which cost an estimated $3-$4.5 million to train.
  • The same recipe produced superhuman play in Barrage Stratego against three world champions, plus state-of-the-art results in Hanabi and dou dizhu.

Ataraxos, a new Stratego AI, beat Pim Niemeijer, described in the Nature paper as 'the most decorated human Stratego player of all time,' by 15 wins to 1 with 4 draws across a 20-game match. Training ran on 16 NVIDIA H100 GPUs for a week and cost under $8,000.

The gap to prior work is the headline the authors lead with. DeepNash, the previous Stratego AI, was trained at an estimated $3 to $4.5 million; Ataraxos used roughly one five-hundredth the compute, one thirtieth the self-play games, and one hundredth the training examples, according to the Nature paper from researchers at Carnegie Mellon, NYU, Stanford and MIT.

The same recipe produced a superhuman AI for Barrage Stratego, beating three multi-time world champions across four 50-game series, and state-of-the-art results on Hanabi and dou dizhu.

The authors frame this as more than a Stratego result. 'The success of this approach across adversarial, cooperative and team games,' they write, 'establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information.' Two researchers we track posted the paper within hours of publication.

The 20-game sample against one human is small, and the paper tests the design pattern only across board and card games. Whether it carries to messier hidden-information domains is a question the paper itself does not pose.

Shared on Bluesky by 2 AI experts