nature.com web signal

Ataraxos beats Stratego champion on $8,000 of compute

TL;DR

  • Ataraxos beat decorated Stratego player Pim Niemeijer 15 wins to 1 loss with 4 draws over 20 games, the first superhuman result reported in the game.
  • Training cost was roughly $8,000, about 1/500th the compute of DeepMind's prior DeepNash effort, which the authors estimate at $3 million to $4.5 million.
  • Beyond Stratego, Ataraxos reached state of the art on Hanabi across 2 to 5 player variants, beat Barrage Stratego world champions, and defeated PerfectDou at Dou Dizhu.

A system called Ataraxos beat Pim Niemeijer, described by the paper as the most decorated Stratego player historically, by 15 wins to 1 loss with 4 draws across a 20-game series. The paper in Nature reports training cost of roughly $8,000, about 1/500th the compute of DeepMind's earlier DeepNash effort, which the authors estimate at $3 million to $4.5 million.

The work is by Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter and Gabriele Farina, researchers at Carnegie Mellon, NYU, Stanford and MIT. Beyond Stratego, Ataraxos beat three multi-time world champions at Barrage Stratego, reached state of the art on Hanabi across 2 to 5 player variants, and beat PerfectDou on Dou Dizhu. Two of the researchers we follow in our Who's Who had posted the link.

The scale is the point. "Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another," the authors write. Stratego carries over 10^33 possible piece configurations; Hanabi exceeds 5 x 10^27 possible deals; Dou Dizhu over 3 x 10^24.

Training used 16 NVIDIA H100 GPUs for a week on the reinforcement-learning side plus 4 H100s for four days on a separate belief network that predicts opponent pieces. Ataraxos trained on 163 million games, against DeepNash's 5.5 billion. The authors' claim is that "reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information."

Shared on Bluesky by 2 AI experts