Ataraxos AI beats Stratego champion 15-1-4 in Nature paper
TL;DR
- Ataraxos, an academic AI from CMU, MIT, NYU and Stanford, defeated Stratego world champion Pim Niemeijer 15-1-4, the first superhuman result in the game.
- Training cost was roughly $8,000, using 1/30th the games and 1/100th the training examples of the predecessor system DeepNash.
- The same technique produced a superhuman Barrage Stratego AI and state-of-the-art AIs for Hanabi and dou dizhu.
Ataraxos, an academic AI system built by researchers at Carnegie Mellon, MIT, NYU and Stanford, defeated the "most decorated human Stratego player of all time" 15-1-4 after training for roughly $8,000, according to a paper published in Nature on 30 September.
Stratego has been imperfect-information AI's standing holdout. The abstract puts the problem bluntly: "the presence of large amounts of hidden information renders established reinforcement learning and search approaches ineffective," and "even with multimillion-dollar industrial research efforts, top-human-level play at Stratego...has remained beyond the reach of artificial intelligence." The authors call the match against champion Pim Niemeijer "to our knowledge, the first superhuman result in the game's history," achieved "while consuming orders of magnitude less compute and data than previous efforts." They report using roughly 1/30th the games and 1/100th the training examples of the predecessor system DeepNash.
The paper argues the same recipe scales. The authors describe Ataraxos as "based on general techniques that we developed for both self-play reinforcement learning and test-time search under hidden information," and report a superhuman AI for Barrage Stratego alongside state-of-the-art AIs for the cooperative card game Hanabi and the Chinese team-based card game dou dizhu. That cross-game breadth is what carries the authors' bigger claim, that the method "establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information."
Lead authors Samuel Sokota and J. Zico Kolter are at CMU, with co-authors Eugene Vinitsky at NYU, Hengyuan Hu at Stanford, and Zhiyuan Fan and Gabriele Farina at MIT. Two researchers on our radar circulated the paper the day it appeared.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature