nature.com web signal

Ataraxos AI beats Stratego world champion in Nature paper

TL;DR

  • Ataraxos beat 4-time Stratego world champion Pim Niemeijer 15 wins to 1 loss with 4 draws across a 20-game series.
  • Training ran on 16 NVIDIA H100 GPUs for 1 week at a cost of a few thousand dollars, roughly 1/500th of DeepNash's compute.
  • The same recipe set state-of-the-art on Hanabi across 2-5 players, beat PerfectDou and DouZero at Dou Dizhu, and reached superhuman Barrage Stratego.

Ataraxos, an AI developed by researchers at Carnegie Mellon, MIT, Stanford and NYU, beat 4-time Stratego world champion Pim Niemeijer across a 20-game series by 15 wins, 1 loss and 4 draws. The result appears in a paper published in Nature on September 30, 2026.

The compute budget is the striking part. Training ran on 16 NVIDIA H100 graphics processing units for 1 week, which the paper puts at "a few thousand dollars," against DeepMind's prior DeepNash effort that reportedly cost roughly $3-4.5 million. The authors put the gap at "roughly 1/500th of the compute cost."

Stratego has over 10³³ possible piece configurations and keeps each player's pieces hidden from the opponent until contact. The paper frames the broader problem: "Real-world decision-making generally involves hidden information, that is, information that is unknown to one agent but possessed by another."

The architecture combines policy-value networks trained through self-play, a belief network modeling hidden information, and a test-time search that refines decisions using belief sampling. Human opponents reported that Ataraxos "excels relative to humans at long-term positional play, punishing mistakes, defending, playing from an information deficit" and "feels preternaturally lucky, always seeming to have the pieces it needs in the right places."

The approach is not Stratego-specific. The authors also report the first superhuman result on Barrage Stratego against three multi-time world champions, state-of-the-art Hanabi scores across 2-5 player variants (24.654 and 77.53% perfect games in the 2-player setting), and wins over PerfectDou and DouZero on the Chinese card game Dou Dizhu. "The success of this approach across adversarial, cooperative and team games establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information," the paper concludes. Code is public at github.com/AtaraxosAI, and two researchers we track flagged the link within the day.

Shared on Bluesky by 2 AI experts