Ataraxos AI tops Stratego champion at 1/500 compute cost
TL;DR
- Ataraxos beat Pim Niemeijer, the most decorated Stratego player, 15 wins to 1 loss with 4 draws over a 20-game match.
- The system trained on 16 NVIDIA H100 GPUs for one week, roughly 1/500th of the compute cost of prior industrial Stratego efforts.
- The same framework reached new state of the art on cooperative Hanabi and beat PerfectDou and DouZero on Dou Dizhu.
Ataraxos won 15 games, drew 4, and lost 1 to Pim Niemeijer, the most decorated player in Stratego's history, across a 20-game match. The Nature paper published September 30 reports an "85% effective win rate" at the highest level of human play, with an overall championship record of 39-2 against top humans.
The compute is the part that surprised people. The system, built by a team from Carnegie Mellon, MIT, NYU and Stanford, trained on 16 NVIDIA H100 GPUs for 1 week, which the authors describe as "a few thousand dollars," against the millions spent on prior industrial Stratego efforts. Co-author Gabriele Farina told MIT News the system "reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples."
Stratego has roughly 10^66 possible piece configurations and both players start blind to the opponent's setup. "With Stratego, there is an explosion of possible universes you might have to deal with," Farina said. The paper combines a transformer policy-value network, a belief network that models hidden information, and test-time search using update-equivalence methods.
Lead author Samuel Sokota framed the bluffing dynamic plainly. "The more you bluff, the more your opponent expects it, and the less each bluff is worth. It's very different from a setting like chess, where the best move is still the best move."
The same framework reached a new state of the art on cooperative Hanabi and beat PerfectDou and DouZero at the Chinese card game Dou Dizhu. Two researchers we follow had the paper circulating in our tracker on publication day.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature