news.mit.edu web signal

Ataraxos tops world Stratego player 15-1-4 in Nature paper

ai-research

TL;DR

  • Ataraxos beat the world's strongest Stratego player 15-1-4 and posted a 39-2 record at the Stratego world championship.
  • The system reached higher playing strength than DeepMind's DeepNash using under 1/100 of the training examples and under 1/30 of the self-play games.
  • Ataraxos pairs self-play reinforcement learning with decision-time planning and a generative model that estimates hidden opponent pieces.

Ataraxos, a Stratego-playing system built by researchers at MIT, Carnegie Mellon, NYU and Stanford, beat the world's strongest Stratego player 15-1-4 and posted a 39-2 record at the Stratego world championship, according to MIT News. The underlying paper, "Scalable decision-making for games of imperfect information," was published in Nature on September 30, 2026.

The headline is not just that it won, but how cheaply. "Our system reaches strictly higher playing strength than DeepNash (DeepMind's system) while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency," says Gabriele Farina, an MIT assistant professor in electrical engineering and computer science and a principal investigator at the Laboratory for Information and Decision Systems.

Stratego hides each side's pieces until they engage, which balloons the number of possible board states a planner has to reason over. "With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting," Farina says. Ataraxos pairs self-play reinforcement learning with decision-time planning and a generative model that estimates where the opponent's hidden pieces actually are.

"Rather than just guessing blindly, we use decision-time planning to find the most plausible state of the board. Using this generative model allows us to really zoom in on the specific board and opponent we are facing," Farina says. Lead author Samuel Sokota, a graduate student at Carnegie Mellon, frames the challenge by contrast: "It's very different from a setting like chess, where the best move is still the best move no matter how often you've played it."

Farina ties the result back to tasks outside the board. "In the kind of imperfect information tasks you would face in reality, you often don't have the luxury of enumerating through all the possibilities. There are just too many."