nature.com web signal

Ataraxos defeats Stratego champion on $8,000 of compute

TL;DR

  • Ataraxos beat top Stratego player Pim Niemeijer 15 wins to 1, with 4 draws over 20 games, which the authors call the first superhuman result in Stratego's history.
  • The system was trained for roughly $8,000 on 16 NVIDIA H100 GPUs for a week, against an estimated $3-4.5 million for the prior DeepNash effort.
  • The same approach set a new state of the art in cooperative Hanabi across 2-5 player variants and defeated PerfectDou and DouZero at Dou Dizhu.

Ataraxos, an AI system built by researchers at Carnegie Mellon, MIT, NYU and Stanford, defeated the most decorated Stratego player in history with 15 wins, one loss and four draws across 20 games, which the authors call "the first superhuman result in the game's history." The paper in Nature puts the training bill at roughly $8,000, against an estimated $3-4.5 million for the prior frontier system, DeepNash.

The margin matters because of what Stratego actually asks of a learner: hidden piece identities, long horizons, and no cheap enumeration of opponent possibilities. Ataraxos pairs a transformer-based policy-value network with a belief network that samples plausible hidden states, then runs a test-time search to refine moves. The authors describe the core trick as "a coordination between regularization strength and policy update size," with heavy regularization and aggressive updates early in training, giving way to weaker regularization and smaller updates later.

Compute figures carry the story. The team reports 1/500th the compute cost of prior industrial efforts, 1/30th the self-play games, and 1/100th the training examples, with training done on 16 NVIDIA H100 GPUs for one week plus four more H100s for four days on the belief network. A custom CUDA-accelerated Stratego simulator pushed roughly 10 million board state updates per second across 163 million finished games and 208 billion environment steps.

Vincent de Boer, a three-time world champion who played the system in a separate evaluation, called the match setup "a large handicap" for the human side because "Ataraxos could not adapt to Pim" while Pim Niemeijer could study its play across games. Beyond Stratego, the authors report a new state of the art in cooperative Hanabi across 2-5 player variants and wins over PerfectDou and DouZero in the Chinese card game Dou Dizhu, which they frame as a general design pattern for reinforcement learning under large amounts of hidden information.

Shared on Bluesky by 2 AI experts