nature.com web signal

Ataraxos beats Stratego champion at 1/500th the compute cost

TL;DR

  • Ataraxos defeated Pim Niemeijer, the highest-ranked Stratego player of all time, 15-1 with 4 draws across 20 games.
  • The system was trained for under US$8,000 at 2025 prices, roughly 1/500th of DeepMind's earlier DeepNash budget.
  • The same recipe reached superhuman play in Barrage Stratego, Hanabi, and dou dizhu, spanning adversarial, cooperative, and team games.

Ataraxos, a reinforcement-learning system built by researchers at Carnegie Mellon, MIT, Stanford, and New York University, defeated Pim Niemeijer, described as the most decorated Stratego player ever, by 15 wins to 1 across 20 games with four draws. The system was trained for less than US$8,000 at 2025 prices, which the authors put at roughly one five-hundredth of the US$3,000,000 to US$4,500,000 cost of DeepMind's earlier DeepNash system. The result appears in a Nature paper published on 30 September 2026.

The same recipe reached what the authors call superhuman play in three further hidden-information games. In Barrage Stratego, Ataraxos beat three of the four top-ranked players across four 50-game series, the first superhuman result for that variant. In Hanabi, its perfect-game rate ran from 58.05% in the five-player version to 89.90% in the three-player version. In dou dizhu, it set new state-of-the-art role-averaged scores against both DouZero and PerfectDou, the two leading prior systems.

The authors write that the work "establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum." The hidden-information spaces they attack are large: the paper counts over 10^33 possible Stratego piece configurations, over 5 × 10^27 Hanabi deals, and over 3 × 10^24 dou dizhu deals.

Technically, the move is a dynamically damped self-play procedure paired with a belief network that predicts opponents' concealed information, feeding a custom CUDA simulator the paper calls StrategoRolloutBuffer. Training ran for one week on 16 Nvidia H100 GPUs plus four days on four more H100s for the belief network, producing 163 million finished games and 208 billion environment steps for the Stratego run alone. Two of the AI researchers we track in Who's Who had posted the paper to their networks.

The headline cost comparison rests on the Ataraxos authors' own figures for the earlier DeepNash budget. The paper reports P values below 2.6 × 10^-4 for the Stratego head-to-head and below 1.3 × 10^-5 for the Barrage series.

Shared on Bluesky by 2 AI experts