Ataraxos beats Stratego champion Niemeijer 15-1-4 in Nature
TL;DR
- Ataraxos went 15-1-4 across a 20-game series against four-time world champion Pim Niemeijer, a margin without precedent at Stratego's top level.
- The system used roughly 1/500th the compute, 1/30th the self-play games, and 1/100th the training examples of DeepMind's earlier DeepNash.
- The same recipe reports first superhuman results on Barrage Stratego and new state-of-the-art on Hanabi (2 to 5 players) and Dou Dizhu.
Ataraxos, an AI system built by researchers at Carnegie Mellon, MIT, Stanford and NYU, went 15 wins, 1 loss and 4 draws across a 20-game series against Pim Niemeijer, holder of four Stratego world championships and more than 600 weeks ranked first. Across the broader championship test the system finished 39-2. The result, published in Nature on September 30, is a margin without precedent at the top level of Stratego play.
Stratego hides most of its state from both players, with piece identities concealed behind their backs. The researchers put the number of possible piece configurations at over 10^66. "With Stratego, there is an explosion of possible universes you might have to deal with," MIT's Gabriele Farina told the MIT press office. "Rather than just guessing blindly, we use decision-time planning to find the most plausible state of the board."
The headline isn't really the match; it's the budget. Ataraxos used roughly 1/500th the compute cost, 1/30th the self-play games, and 1/100th the training examples of DeepMind's earlier DeepNash. The paper puts the cost at a few thousand dollars against what it describes as previous industrial efforts spanning millions. The authors credit a method they call dynamic damping, which pairs strong regularization with aggressive policy updates early in training and inverts both as the agent sharpens. Samuel Sokota of Carnegie Mellon frames the logic plainly: "The more you bluff, the more your opponent expects it, and the less each bluff is worth."
The same recipe carries beyond Stratego. The team reports first superhuman results on Barrage Stratego, new state-of-the-art on Hanabi across its 2 to 5 player variants, and a decisive win over previous Dou Dizhu bots. The abstract concludes that "reinforcement learning and search are no longer precluded from high performance by the presence of large amounts of hidden information." Two researchers we track posted the paper the day it went up.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature