nature.com web signal

Ataraxos beats Stratego world champion on a few-thousand-dollar budget

TL;DR

  • Ataraxos beat Pim Niemeijer, the paper's 'most decorated Stratego player historically,' 15 wins to 1 loss with 4 draws.
  • The system trained on ~160 million games versus DeepNash's ~5.5 billion, roughly 1/500th the compute.
  • The same recipe set state-of-the-art on Hanabi across 2- to 5-player variants and beat prior best bots at dou dizhu.

Ataraxos, an AI system built by researchers at Carnegie Mellon, NYU, Stanford and MIT, beat Pim Niemeijer, described in the paper as the most decorated Stratego player historically, 15 wins to 1 loss with 4 draws. The authors put the total training cost at "a few thousand dollars."

The result, published in Nature, takes on a longstanding wall in game AI. Reinforcement learning and search work spectacularly on chess and Go but stall when players cannot see each other's pieces. "Hidden information renders established reinforcement learning and search approaches ineffective," the authors write. Stratego, with its concealed setup, has been the canonical hard case.

DeepMind's prior DeepNash system cleared the same bar at very different scale. Ataraxos trained on roughly 160 million games and about 50 billion examples; DeepNash used roughly 5.5 billion games and 5 to 10 trillion examples. The authors put the compute ratio at about 1/500th. Much of the gap runs through a CUDA-accelerated simulator the team clocks at "roughly 10 million board state updates per second."

The system pairs transformer-based policy-value networks with a belief network that models the opponent's hidden pieces, plus a test-time search that performs "single damped reinforcement learning update steps at decision points." Training uses what the paper calls "dynamically damped self-play" to suppress the "cyclical, divergent or chaotic learning dynamics" that plague self-play in imperfect-information games.

Beyond Stratego, the same recipe produced the first superhuman result for Barrage Stratego, set state-of-the-art scores on Hanabi across 2- to 5-player variants, and beat prior best bots at dou dizhu. In a 40-game demo at the Stratego World Championship, the authors log a 95% effective win rate.

Two researchers we track in our Who's Who directory posted the Nature link the day it went up.

Shared on Bluesky by 2 AI experts