Ataraxos AI tops Stratego champion on 1/500th the compute
TL;DR
- Ataraxos defeated Pim Niemeijer, the most decorated Stratego player, 15-1-4, and finished 39-2 against top championship players overall.
- Training used less than 1/100th of the examples and 1/30th of the self-play games of DeepMind's DeepNash, costing thousands rather than millions.
- The same framework produced the first superhuman Barrage Stratego AI and state-of-the-art results in Hanabi and Dou Dizhu.
Ataraxos, an AI built by researchers at Carnegie Mellon, MIT, Stanford and NYU, beat Pim Niemeijer, the most decorated Stratego player, 15 wins to 1 loss with 4 draws, and finished 39-2 against top championship players overall. The paper, published in Nature on 30 September, describes the first superhuman result in Stratego's history, in a game with more than 10^66 possible piece configurations.
The cost is the headline. The authors say Ataraxos reached strictly higher playing strength than DeepMind's DeepNash, the previous Stratego benchmark, while using less than one hundredth of the training examples and less than one thirtieth of the self-play games. Training ran into the thousands of dollars rather than the millions.
The system combines self-play reinforcement learning with test-time search and a belief network that models what the opponent might be hiding. "With Stratego, there is an explosion of possible universes you might have to deal with," said Gabriele Farina of MIT, the senior author. "Rather than just guessing blindly, we use decision-time planning to find the most plausible state of the board." Lead author Samuel Sokota of Carnegie Mellon described the bluffing dynamic the agent learned to meter: "The more you bluff, the more your opponent expects it, and the less each bluff is worth."
The same framework beat three multi-time world champions in Barrage Stratego, posted state-of-the-art results in Hanabi across its 2-to-5-player variants, and outperformed the earlier PerfectDou and DouZero bots in Dou Dizhu. Two researchers we track flagged the paper the day it published.
The authors suggest the technique could carry over to military maneuvers, business negotiations and cybersecurity. The paper itself reports no results in those domains.
Shared on Bluesky by 2 AI experts
-
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. www.nature.com/articles/s41...
View on Bluesky →
Originally reported by nature.com
Read the original article →Original headline: Scalable decision-making for games of imperfect information - Nature