nature.com web signal

Ataraxos defeats top Stratego champion on sub-$8K budget

TL;DR

  • Ataraxos beat Pim Niemeijer, called the most decorated human Stratego player, 15 wins to 1 loss with 4 draws across 20 games.
  • Training ran on 16 NVIDIA H100 GPUs for one week at under US$8,000, roughly 1/500th the compute of DeepMind's earlier DeepNash effort.
  • The same method produced a superhuman Barrage Stratego AI and state-of-the-art results in Hanabi and dou dizhu.

Ataraxos, a new AI for the hidden-information board wargame Stratego, defeated Pim Niemeijer — described by the authors as the most decorated human Stratego player of all time — 15 wins, 1 loss and 4 draws across a 20-game series, in what its authors call "the first superhuman result in the game's history." Training cost less than US$8,000 at 2025 pricing on 16 NVIDIA H100 GPUs for one week, roughly 1/500th the compute DeepMind spent on its earlier DeepNash effort (US$3–4.5 million).

The system is described in Nature by a group spanning Carnegie Mellon, NYU, Stanford and MIT, with Samuel Sokota of CMU's Machine Learning Department as first author. Alongside the Niemeijer series the authors ran a demonstration match at the 2025 World Championship, recording 38 wins to 2 losses across 40 games, and against top-ranked two-time world champions of the Barrage Stratego variant Ataraxos won all four 50-game series.

The paper's qualitative read on how the AI plays is unusually direct: "Ataraxos feels preternaturally lucky, always seeming to have the pieces it needs in the right places." Its edge over humans, the authors write, lies in "long-term positional play, punishing mistakes, defending, playing from an information deficit, transitioning from middle- to end-game."

The same algorithmic recipe — self-play reinforcement learning with a belief network, test-time search, and a coordinated schedule for regularization strength and policy update size — produced state-of-the-art results in the cooperative card game Hanabi (77.53% perfect games in the 2-player setting, scoring 24.654 out of 25) and in dou dizhu, where Ataraxos outscored PerfectDou by +0.199 and DouZero by +0.350 per role-averaged game.

The authors frame it as a general design pattern: "The success of these techniques across adversarial, cooperative and team games shows that reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information." Niemeijer was paid a US$1,000 base plus US$100 per win and US$50 per draw under a CMU IRB-approved protocol.

Shared on Bluesky by 2 AI experts