nature.com web signal

Ataraxos AI beats top Stratego player at 1/500th the cost

TL;DR

  • Ataraxos beat Pim Niemeijer, the most decorated Stratego player in the game's history, 15 wins to 1 loss with 4 draws over 20 games.
  • Training ran roughly $8,000 on 16 NVIDIA H100s for a week, about 1/500th the compute cost of DeepMind's prior DeepNash system.
  • The same recipe set new state-of-the-art marks on Hanabi for 2 to 5 players, Barrage Stratego, and Dou Dizhu.

Ataraxos, a reinforcement-learning agent built by researchers at Carnegie Mellon, MIT, Stanford and NYU, beat Pim Niemeijer, the most decorated Stratego player in the game's history, 15 wins to 1 loss with 4 draws over a 20-game series. The result, published in Nature on September 30, is described by the authors as the first superhuman result in Stratego, a two-player game where each side keeps its pieces hidden from the opponent.

The second headline is the price. Training Ataraxos ran roughly $8,000 on 16 NVIDIA H100s for one week, plus four more H100s for four days to train its belief network. DeepMind's prior Stratego system, DeepNash, cost between $3 million and $4.5 million. The paper reports "1/500th the compute cost and 1/30th the self-play games."

The same recipe, a policy-value network paired with a belief network modelling hidden information and a test-time search procedure, set new state-of-the-art marks across Hanabi for 2 to 5 players, Barrage Stratego, where it defeated three world champions in 50-game series, and Dou Dizhu, where it beat the previous leader PerfectDou. In a demonstration round at a World Championship, Ataraxos went 38 wins to 2 losses with no draws.

The core trick is a schedule the authors call dynamic regularization coupling: "stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller policy updates late in self-play." That shift, they argue, is what makes self-play tractable when much of the state is hidden from the learner.

Ataraxos trained on 163 million finished games and 208 billion environment steps, running on a custom GPU-accelerated simulator the authors clock at around 10 million board state updates per second. Two researchers in our tracking network posted the paper on publication day.

Shared on Bluesky by 2 AI experts