nature.com web signal

Ataraxos beats top Stratego player on $8,000 compute budget

TL;DR

  • Ataraxos defeated the top-ranked Stratego player Pim Niemeijer with 15 wins, 1 loss and 4 draws, an 85% effective win rate.
  • The system trained for under $8,000, roughly 1/500th the compute of DeepMind's earlier DeepNash, which cost about $3 to $4.5 million.
  • The same approach set state-of-the-art Hanabi scores, beat three Barrage Stratego world champions, and defeated PerfectDou and DouZero at dou dizhu.

Ataraxos, an AI system built by researchers at Carnegie Mellon, MIT, NYU and Stanford, defeated the most decorated competitive Stratego player, Pim Niemeijer, 15 games to 1 with 4 draws. The result, published in Nature, is described in the paper as a "superhuman result in the game's history".

What stands out alongside the win is the budget. The authors say Ataraxos was trained for under $8,000, against roughly $3 to $4.5 million for DeepMind's earlier DeepNash system, consuming 1/500th the compute, 1/30th the games, and 1/100th the training examples of that prior effort. The run used 16 NVIDIA H100 GPUs for one week to train the policy-value network, plus four H100s for four days for a belief network, over 163 million finished games.

The same approach beat three multi-time world champions at Barrage Stratego, set new state-of-the-art scores on Hanabi across its 2- to 5-player variants, and defeated the prior best dou dizhu agents, PerfectDou and DouZero.

Ataraxos combines a self-play policy-value network, a belief network that models the hidden pieces, and a test-time search procedure that refines its moves at play time. The team credits a scheduler they call dynamic damping, which applies "stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller policy updates late".

The authors frame imperfect information as "a characterizing feature of real-world settings, including financial markets, military conflict and negotiations", and conclude: "The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning."

Shared on Bluesky by 2 AI experts