Ataraxos tops decorated Stratego champion 15-1-4 in Nature
TL;DR
- Ataraxos beat Pim Niemeijer, called the most decorated Stratego player of all time, with 15 wins, 1 loss and 4 draws across 20 games.
- Training ran at a few thousand dollars and about 160 million games, versus DeepNash's $3,000,000 to $4,500,000 and about 5.5 billion games.
- The same recipe also produced a superhuman Barrage Stratego player and state-of-the-art agents for Hanabi and dou dizhu.
Ataraxos, an AI system built by researchers at MIT, Carnegie Mellon, NYU and Stanford, beat Pim Niemeijer, described in the Nature paper as "the most decorated Stratego player of all time," by 15 wins, 1 loss and 4 draws across a 20-game series. The authors call that an 85% effective win rate.
The compute line is the part that will land in labs. Training ran "at a total cost of a few thousand dollars," against DeepMind's DeepNash figure of "$3,000,000 and $4,500,000." Sample efficiency moves in the same direction: roughly 160 million games of self-play, where DeepNash used "about 5.5 billion games." The paper, lead-authored by Samuel Sokota at Carnegie Mellon with coauthors including J. Zico Kolter and MIT's Gabriele Farina, appears in Nature volume 658, pages 55 to 59, and two researchers we track circulated it the same day.
The humans had opinions. Players reported that the system "feels preternaturally lucky, always seeming to have the pieces it needs in the right places," "takes gambles that humans consider arrogant," and "fights more bitterly than humans when it is behind." They noted a stronger preference than humans for preserving Scouts deep into the game, more willingness to accept a draw, and sparing use of certain bluffs. Human participants received "$1,000 for participating...with an additional $100 for each win."
The architecture is three pieces: a policy-value network trained by self-play, a belief network modelling the hidden pieces, and test-time search. The authors frame the training trick as "coordination between regularization strength and policy update size," with stronger regularization and more aggressive updates early in self-play, and weaker regularization and smaller updates late. They apply the same recipe to Barrage Stratego, Hanabi and dou dizhu.
The paper does not report a direct Ataraxos-versus-DeepNash match; the comparison is on compute and sample budget. The human evaluation is a 20-game series against a single top player at 15+3 time controls, a strong anchor but one opponent.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature