nature.com web signal

Ataraxos beats Stratego champion 15-1-4 for under $8,000

TL;DR

  • Ataraxos beat Pim Niemeijer, the most decorated Stratego player of all time, 15-1-4 in a 20-game series, the first superhuman AI result in Stratego.
  • The training budget was under $8,000 on H100s, versus roughly $3 million to $4.5 million the paper estimates DeepMind's DeepNash would cost today.
  • The same recipe produced a superhuman AI for Barrage Stratego and state-of-the-art results on Hanabi and dou dizhu.

Ataraxos, an AI built by researchers at MIT, Carnegie Mellon, Stanford, and NYU, defeated Pim Niemeijer, described in the paper as "the most decorated human Stratego player of all time," 15 wins to 1 loss with 4 draws in a 20-game series. The paper, published in Nature on September 30, reports this as the first superhuman result in Stratego's history, achieved on a training budget of under $8,000 on H100s, versus the roughly $3 million to $4.5 million the paper estimates DeepMind's prior DeepNash would cost under current pricing.

The approach pairs self-play reinforcement learning with test-time search over hidden information, using a belief network to sample plausible configurations of the opponent's board. "With Stratego, there is an explosion of possible universes you might have to deal with," MIT's Gabriele Farina told MIT News. "Rather than just guessing blindly, we use decision-time planning to find the most plausible state of the board." Farina added that the system "reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples." Two researchers we track shared the paper link within hours of publication.

The same recipe produced a superhuman AI for Barrage Stratego and state-of-the-art results on Hanabi and dou dizhu. The paper frames this as a general template, writing that it "establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information." No results outside those four games appear in the published paper.

Shared on Bluesky by 2 AI experts