nature.com web signal

Ataraxos AI beats world Stratego champion 15-1-4 in Nature

TL;DR

  • Ataraxos beat Pim Niemeijer, Stratego's most decorated player, 15-1-4 and logged a 95% win rate across 40 demos at the 2025 World Championship.
  • Training cost about $8,000 on 16 NVIDIA H100 GPUs for one week, a reported ~500x compute cut versus DeepMind's DeepNash.
  • The same policy-value plus belief-network plus test-time-search recipe hit superhuman Barrage Stratego and new state-of-the-art Hanabi and Dou Dizhu.

An AI system called Ataraxos beat Pim Niemeijer, Stratego's most decorated player in history, 15 games to 1 with 4 draws. The compute bill came to roughly $8,000.

Those numbers appear in a paper published in Nature on September 30. The point of pairing them is that DeepMind's DeepNash system, the previous state of the art, reportedly cost somewhere between $3 million and $4.5 million to train. Ataraxos used around 1/500th the compute, 1/30th the self-play games, and 1/100th the training examples. The reinforcement-learning phase ran on 16 NVIDIA H100 GPUs for one week.

Stratego hides nearly everything. Each side's 40 pieces carry identities the opponent only learns when two pieces collide, which leaves over 10^66 possible piece configurations. "With Stratego, there is an explosion of possible universes you might have to deal with," MIT co-author Gabriele Farina told MIT News. Across 40 demonstration games at the 2025 World Championship, the system logged a 95% effective win rate; MIT reports a 39-2 record against top human players at the event.

The system has three parts: a policy-value network trained via self-play, a belief network that models what the opponent might be holding, and a test-time search that samples possible worlds and runs further RL updates against them. The authors call the training recipe "dynamically damped" self-play. The same approach produced superhuman play on Barrage Stratego against three world champions, new state-of-the-art Hanabi across the 2- to 5-player variants, and bots that outperform PerfectDou and DouZero at Dou Dizhu.

The byline runs across CMU (Samuel Sokota and J. Zico Kolter), MIT (Zhiyuan Fan, Farina), Stanford (Hengyuan Hu) and NYU (Eugene Vinitsky), and two researchers on our Who's Who radar had already circulated the link. Sokota flagged the twist chess intuition misses: "The more you bluff, the more your opponent expects it, and the less each bluff is worth."

Shared on Bluesky by 2 AI experts