Ataraxos beats top Stratego human 15-1-4 at 1/500th cost
TL;DR
- Ataraxos defeated Pim Niemeijer, Stratego's most decorated player, 15 wins to 1 loss with 4 draws, and went 39-2 against top human opposition overall.
- The system trained for roughly $8,000, versus the $3 million to $4.5 million spent on DeepMind's prior superhuman Stratego agent DeepNash.
- The same framework produced the first superhuman play in Barrage Stratego, new state-of-the-art in Hanabi across 2 to 5 players, and wins over PerfectDou and DouZero in Dou Dizhu.
An academic AI called Ataraxos beat Pim Niemeijer, the most decorated Stratego player on record, 15 wins to 1 loss with 4 draws, and did it on a training budget of roughly $8,000.
That is the result reported in a Nature paper from Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter and Gabriele Farina, split across CMU, NYU, Stanford and MIT. Stratego, where neither player sees the other's pieces until contact, has more than 10^66 possible piece configurations, exponentially more than chess. The prior superhuman effort on it, DeepMind's DeepNash, cost somewhere between $3 million and $4.5 million to train.
"With Stratego, there is an explosion of possible universes you might have to deal with," Farina told MIT News. Ataraxos pairs a policy-value network trained via self-play with a separate belief network that models the opponent's hidden pieces, then refines its moves at test time with search. "Our system reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples," Farina said. The Nature paper also reports less than 1/30th of the self-play games and roughly 1/500th of the compute of the prior effort.
Against humans the system went 39-2 across top-player opposition before taking the Niemeijer series. "Ataraxos is good at calculating risk in a way that humans are not," Farina said. Sokota described one of the training dynamics plainly: "The more you bluff, the more your opponent expects it, and the less each bluff is worth."
The same framework produced the first superhuman results in Barrage Stratego, where it defeated three multi-time world champions; new state-of-the-art play in Hanabi across its 2- to 5-player variants, with the 2-player result arriving on two orders of magnitude less compute than previous methods; and wins over PerfectDou and DouZero in the Chinese card game Dou Dizhu. Two of the researchers we track in our Who's Who directory were already passing the paper around within hours of its release.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature