nature.com web signal

Ataraxos beats top Stratego champion for under $8,000

TL;DR

  • Ataraxos won 15 of 20 games against the player the authors call the most decorated human Stratego player of all time, with one loss and four draws.
  • The paper reports training costs of fewer than $8,000 at 2025 pricing, roughly one five-hundredth the compute of prior industrial efforts.
  • The same technique produced state-of-the-art results in Barrage Stratego, Hanabi and dou dizhu, spanning adversarial, cooperative and team games.

A team from Carnegie Mellon, MIT, NYU and Stanford reports in Nature that their system, Ataraxos, beat Pim Niemeijer — described in the paper as "the most decorated human Stratego player of all time" — by 15 wins to 1 across a 20-game series, with four draws. The authors put the training bill at fewer than $8,000 at 2025 pricing.

Stratego has been the field's standing embarrassment for games of large-scale hidden information. "Even with multimillion-dollar industrial research efforts, top-human-level play at Stratego — a board wargame with hidden information on a massive scale — has remained beyond the reach of artificial intelligence," the abstract says. Ataraxos is pitched as the first superhuman result in the game's history, consuming "orders of magnitude less compute and data than previous efforts."

The authors — Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter and Gabriele Farina — report training in roughly one week on 16 NVIDIA H100 GPUs, playing 163 million games across 208 billion environment steps. They quote a compute cost of about one five-hundredth of the earlier DeepNash effort, with about one-thirtieth the self-play games and one-hundredth the training examples.

The same recipe produced state-of-the-art results in three other hidden-information games: Barrage Stratego against world champions across four 50-game series, Hanabi in its 2-to-5 player variants, and the Chinese card game dou dizhu. The authors describe the result as establishing "a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making." Two researchers we track posted the paper when it appeared.

Shared on Bluesky by 2 AI experts