The Unbothered Machine: Ataraxos Crushes World-Class Stratego Players at 1% of DeepMind's Training Cost
Researchers from MIT, CMU, NYU and Stanford published Ataraxos in Nature — a Stratego AI that beat the world champion 15-1-4 while using less than 1/100th of DeepNash's training data and 16 H100 GPUs for a single week.
On September 30, 2026, a paper appeared in Nature with an unusual protagonist: a system named after a Greek word for someone who is unbothered — free from anxiety. The name turned out to be apt. Ataraxos, an AI developed by researchers from MIT, Carnegie Mellon University, New York University, and Stanford University, walked into the Stratego world championship, posted a 39-2 record against top human players, and defeated the strongest Stratego player on Earth by a record margin of 15-1-4.
It did so while spending a rounding error of what previous state-of-the-art systems burned in training.
Why Stratego Is the Hard One
Every few years, an AI conquers a game and the world takes notice. Chess fell, Go fell, poker fell. Stratego, a two-player board wargame often described as military chess, has quietly resisted the pattern — and the reason is structural, not incidental.
Stratego is a game of imperfect information. Each player arranges 40 pieces on their side of the board, and the identity of every opposing piece remains hidden until two pieces collide, at which point the lower-ranked piece is eliminated. The number of possible piece configurations exceeds 10^66 — an exponentially larger space than chess. You cannot simply enumerate your way to the best move.
“With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting,” says Gabriele Farina, an assistant professor in MIT’s Department of Electrical Engineering and Computer Science (EECS), principal investigator at the Laboratory for Information and Decision Systems (LIDS), and senior author of the paper.
The deeper problem is that hidden information breaks a convenient assumption. “The more you bluff, the more your opponent expects it, and the less each bluff is worth. It’s not obvious how to reason about that,” explains lead author Samuel Sokota, a graduate student at Carnegie Mellon. “It’s very different from a setting like chess, where the best move is still the best move no matter how often you’ve played it.”
Google DeepMind’s DeepNash was the previous high-water mark, but even with sophisticated, computationally demanding operations behind it, the system was never able to definitively beat top human Stratego players.
A Two-Pronged Approach
Ataraxos combines two techniques in a way the researchers say was the missing piece for superhuman play.
First, a blueprint strategy via efficient self-play. The model plays against itself millions of times to learn a strong baseline strategy. The team designed especially efficient training algorithms that let Ataraxos learn dramatically faster than prior methods while ensuring it did not get stuck trying to predict every possible move branch. This is where the cost savings live: “Our system reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency,” Farina says.
Training ran on 16 H100 GPUs for one week — a budget accessible to a university lab rather than a hyperscaler.
Second, decision-time planning with a generative model. During an actual game, the blueprint serves only as the starting point. Before acting, Ataraxos refines its choices on the fly: a generative model estimates the probable identities of the opponent’s hidden pieces, and the system evaluates future choices against that estimate before committing to a move. “Rather than just guessing blindly, we use decision-time planning to find the most plausible state of the board. Using this generative model allows us to really zoom in on the specific board and opponent we are facing,” Farina says.
That on-the-fly opponent modeling is what earned the system its name. “Ataraxos is good at calculating risk in a way that humans are not. A human might start freaking out if their most valuable piece is exposed, but the bot can be surprisingly composed. It doesn’t overcorrect and give away its secrets,” Farina notes.
It Generalizes
To test whether the method was a Stratego one-off, the researchers adapted Ataraxos to other imperfect-information games with different rules and designs: Barrage Stratego (a faster-paced variant with fewer pieces), Hanabi (a cooperative card game with many players), and Dou Dizhu (a Chinese card game in which two players cooperate against a third). The system achieved superhuman performance in each instance — evidence that the combination of efficient self-play and generative decision-time planning is a general recipe, not a game-specific hack.
Why This Matters Beyond the Board
Imperfect information is not a game-genre curiosity — it is the default condition of the real world. Traders in financial markets do not know the rationale behind others’ trades. Military forces do not have full knowledge of enemy positions. Negotiators, cybersecurity defenders, and auction bidders all reason under partial visibility, where every action leaks information and changes the opponent’s beliefs.
The MIT team is explicit about the ambition: such systems could someday help decision-makers select ideal strategies for military maneuvers, business negotiations, or cybersecurity — domains where “you often don’t have the luxury of enumerating through all the possibilities. There are just too many,” as Farina puts it. “Having AI algorithms that are general purpose and can provably perform this challenging task so well is a big step forward.”
The efficiency angle sharpens the significance. If superhuman play in deeply imperfect-information settings used to cost millions of dollars in training compute, and now costs 16 H100s for a week, the barrier to entry for strategic AI has dropped by orders of magnitude. That democratizes research — and lowers the threshold for adversarial applications alike.
The Caveat: Interpretability Before Adoption
The researchers are careful about what comes next. Before such a system could advise a human in a consequential setting, its reasoning has to be auditable. “Humans must have the final say in whether a recommendation is followed, so before adoption can happen, we need a way to audit the model’s decisions. We still have a long way to go, but I hope these algorithms can be the foundation for a lot more work to come,” Farina says. The team’s stated next step is building interpretability measures into Ataraxos so it can explain its decisions in human-understandable terms.
The paper, “Scalable decision-making for games of imperfect information,” is authored by Samuel Sokota (CMU), Eugene Vinitsky (NYU), Zico Kolter (CMU), Hengyuan Hu (Stanford), Zhiyuan Fan (MIT), and Gabriele Farina (MIT). The research was funded in part by the Office of Naval Research, the National Science Foundation, NYU’s C2SMART Center, and a Schmidt Sciences AI2050 Early Career Fellowship.
A composed machine that never panics, models what you’re hiding, and plans accordingly — priced at a university budget. The unbothered champion may be the most interesting AI story of the week that isn’t about a frontier lab.