news.mit.edu web signal

MIT-Led Ataraxos Beats World Stratego Champion 15-1-4 Using 1% of DeepNash's Training Examples

ai-research

Summary

Researchers from MIT, CMU, NYU and Stanford unveiled Ataraxos, a Stratego-playing system that defeated the world's top human player 15-1-4 and went 39-2 at the world championship, publishing results in Nature. Ataraxos combines self-play reinforcement learning with decision-time planning and a generative model that estimates hidden opponent pieces, reaching higher strength than DeepMind's DeepNash while using less than 1/100 of the training examples and 1/30 of the self-play games. Training ran on 16 H100 GPUs for one week.