nature.com web signal

Ataraxos AI beats Stratego world champion for under $8,000

TL;DR

  • Ataraxos, from CMU, MIT, NYU and Stanford, beat four-time Stratego world champion Pim Niemeijer 15-1-4 across 20 games.
  • Training cost was under US$8,000 against DeepMind DeepNash's roughly US$3M-$4.5M, with 160 million self-play games versus 5.5 billion.
  • The same approach set new state of the art on Hanabi and Dou Dizhu and reached superhuman play in Barrage Stratego.

An academic team built an AI that beat the world's most decorated Stratego player 15 wins to 1 for a training cost under $8,000. The system, called Ataraxos (Greek for "free from anxiety"), is described in a Nature paper by researchers at Carnegie Mellon, MIT, NYU and Stanford.

The reference point is DeepMind's DeepNash, which reached expert-level Stratego after roughly 5.5 billion self-play games at a cost the paper puts between US$3,000,000 and US$4,500,000. Ataraxos used about 160 million games and about 50 billion training examples.

"Our system reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples," Gabriele Farina, one of the lead authors, said in an MIT News release.

Stratego hides piece identities until contact, which is why it had resisted the methods that cracked chess and Go. "With Stratego, there is an explosion of possible universes you might have to deal with," Farina said; the paper cites more than 10^66 possible piece arrangements. Samuel Sokota, the paper's first author and a CMU graduate student, framed the problem this way: "It's very different from a setting like chess, where the best move is still the best move."

Against four-time world champion Pim Niemeijer, Ataraxos went 15-1-4 across 20 games. At the Stratego world championship, it finished 39-2 against top human players. The paper also reports superhuman performance in Barrage Stratego, "new state of the art with statistical significance for all variants" of Hanabi, and gains over previous bests in the Chinese card game Dou Dizhu.

The method stacks three components: a policy-value network trained by self-play, a belief network that models what the opponent might be holding, and a test-time search that samples possible world states before each move. "Ataraxos is good at calculating risk in a way that humans are not," Farina said.

The paper does not claim transfer beyond games.

Shared on Bluesky by 2 AI experts