← Back to all posts
News

Ataraxos Beats Stratego's Best Human 15-1-4 in Nature Paper, Trained for Under $8,000

September 30, 2026 · 17:08 UTC · News
Ataraxos Beats Stratego's Best Human 15-1-4 in Nature Paper, Trained for Under $8,000

TL;DR

Stratego, the hidden-piece board wargame that DeepMind spent millions on without reaching top human level, now has a clearly superhuman AI. Ataraxos, published in Nature today by researchers from Carnegie Mellon, MIT, NYU and Stanford, beat Pim Niemeijer, the most decorated Stratego player ever, 15 wins, 1 loss and 4 draws over 20 games. Its main training run used 16 H100s for one week, which the authors put at under $8,000. The same recipe set new records in Hanabi and dou dizhu, and the code is public under an MIT license. If you build agents that act without seeing the whole board, this is the paper to read this week.


What happened

The paper, "Scalable decision-making for games of imperfect information", went online in Nature on September 30, 2026. Authors are Samuel Sokota (CMU), Eugene Vinitsky (NYU), Hengyuan Hu (Stanford), Zhiyuan Fan (MIT), J. Zico Kolter (CMU) and Gabriele Farina (MIT). A preprint with the headline Stratego result, "Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search", has been on arXiv since November 2025. Today's version is the peer-reviewed one, with the full method, cost accounting and results on three more games.

The headline match: a 20-game series against Niemeijer, played over three weeks under the default 15+3 Strategus time control. Niemeijer's resume per the paper: 4 world championships, 15 Dutch national titles, 2 online world championships and more than 600 weeks ranked number one. Ataraxos won 15-1-4, an 85% effective win rate (draws count as half). At the top level of Stratego, where margins are usually razor thin, that is a rout.

To keep the human honest, Niemeijer was paid $1,000 to play plus $100 per win and $50 per draw. By that formula, his one win and four draws earned him $300 in bonuses. The machine, per the paper, felt "preternaturally lucky" to play against, which is what a well-calibrated bluffer looks like from the other side of the board.

Ataraxos results vs humans (wins / losses / draws) vs Niemeijer 15 W4 D 20 games WC 2025 demo 38 W 40 games 85%95% win loss draw (% = effective win rate)
One loss in 20 games against the best player in Stratego history, and 38 of 40 against world championship attendees.

A second data point: at the 2025 Stratego World Championship (August 1 to 3), attendees could play Ataraxos. Across 40 games it went 38-2-0, a 95% effective win rate. (MIT's press release says 39-2; the Nature paper says 38 wins, 2 losses, 0 draws in 40 games, and we go with the paper.)

Why Stratego was the holdout

Chess and Go are perfect-information games: the right move is the right move no matter how often you play it. Stratego hides every piece identity until two pieces collide, and the paper counts more than 1033 possible piece configurations. The value of a move depends on what your opponent believes about your pieces, which depends on how you have played so far, which depends on what you think they believe. Poker-style solvers handle this by reasoning over public information, but that cost scales with the amount of hidden information. Texas hold'em has 1,326 possible hands. Stratego is not Texas hold'em.

DeepMind's DeepNash (2022) was the previous best effort. The Nature paper notes it won 42 of 50 counted games on the Gravon site in April 2022, but by then only 25 players were ranked there, and DeepNash lost to most of the highest-ranked players who faced it, including Niemeijer. The Ataraxos team asked DeepMind for a head-to-head match. DeepMind said the DeepNash code "is no longer functional." So the comparison is on resources, and there it is lopsided.

How much less Ataraxos used than DeepNash (x fewer) compute cost~500x train examples~100x self-play games~30x DeepNash est. $3.0M-4.5M (1,024 TPU v3 nodes, 2-3 months) vs <$8K
Stronger play on roughly 1/500th of the compute bill, per the paper's own cost accounting at 2025 prices.

The paper's accounting: DeepNash trained on 1,024 TPU v3 nodes for two to three months, roughly $3.0M to $4.5M at 2025 prices. Ataraxos trained its policy on 16 H100s for one week and its belief model on 4 H100s for four days, under $8,000. DeepNash consumed about 5.5 billion games and 5 to 10 trillion training examples; Ataraxos about 160 million games and 50 billion examples. The authors are careful to point out that fewer games and examples means the savings come from sample efficiency, not only from faster code.

How it works

Ataraxos is three parts that the authors pitch as a reusable design pattern for any imperfect-information problem with a fast simulator:

  • Policy-value networks trained by self-play. Two transformer networks, one for the opening piece set-up and one for moves, each trained by reinforcement learning against itself. A technique the paper calls dynamically damped self-play keeps updates on-policy without either stalling or blowing up.
  • A belief network. Trained on the final policy's self-play games, where the hidden pieces are known, to predict what the opponent is probably holding given what you have seen.
  • Test-time search. Before each move, Ataraxos samples plausible hidden boards from the belief network, looks ahead on them, and refines its move choice.
self-play RL16 H100 x 1 wk belief network4 H100 x 4 days test-time searchsample boards, plan move
Learn a blueprint by self-play, learn to guess hidden pieces, then plan on those guesses at every move.

The belief network is the part that did the heavy lifting. MIT News describes using a generative model for decision-time planning as "the missing piece" that got the system to superhuman play. Think of it like a poker pro who, before calling, mentally deals out the dozen hands the other player most plausibly holds and checks the call against each one, instead of either guessing one hand or trying to enumerate all of them.

The unglamorous half of the win is engineering. To fit on what the authors call "modest academic compute infrastructure," they wrote a Stratego simulator in CUDA C++ targeting roughly 10 million board-state updates per second. The main run covered 163 million finished games and 208 billion environment steps. If your RL project is bottlenecked on environment throughput, that section of the Methods is worth your time on its own.

It generalizes, at least across games

The same recipe, with game-specific networks, produced:

  • Barrage Stratego: wins over three multi-time world champions, which the authors call the first superhuman result for that variant.
  • Hanabi, the standard cooperative imperfect-information benchmark: new state of the art for two to five players. The two-player score is 24.654 out of 25 with 77.53% perfect games, and the paper says it took two orders of magnitude less compute than the previous best.
  • Dou dizhu, the Chinese three-player card game: statistically significant wins over the previous best bots, PerfectDou and DouZero.

That spread matters: adversarial two-player, fully cooperative, and a two-versus-one team game. The claim is a general design pattern, not a Stratego-specific bag of tricks.

What you can actually run

The stratego repo ships the CUDA simulator, self-play RL training and belief-model training, MIT-licensed, with sibling repos for Hanabi and dou dizhu. It targets Linux with CUDA 11.8 and a micromamba environment, and the README's demo config runs training iterations in 10 to 20 seconds on a single RTX A6000. The 20 games against Niemeijer are published as static records at ataraxosai.github.io. The README does not advertise pretrained weights, so plan on training your own.

The caveats

  • No head-to-head with DeepNash. The strength comparison rests on human results and resource counts, because DeepNash can no longer run.
  • A fixed opponent. Niemeijer was told Ataraxos would not adapt to him. The authors argue that not adapting was a handicap for the bot, not the human, but 20 games is still a small sample.
  • Games are not the world. The recipe needs a fast, accurate simulator. MIT's release name-checks military maneuvers, negotiations and cybersecurity, and the work is funded in part by the Office of Naval Research, but none of those domains comes with a CUDA simulator that runs 10 million steps per second.
  • Search has a ceiling. The paper notes its search mimics a single update step, so throwing more test-time compute at it gives bounded gains.

Why builders should care

Most interesting agent problems are imperfect-information problems: pricing against competitors, negotiating, security, any multi-agent setting where the other side hides its hand. The LLM agent world mostly handles this by vibes. Ataraxos is a concrete, cheap and now peer-reviewed template: self-play for a blueprint, a learned belief model over hidden state, and search that samples from it at decision time. The fact that a university team did it for less than the price of a used car, where an industrial lab spent millions and fell short, is the part that should make you reread your own compute budget.

Key Takeaways

  • Ataraxos beat Stratego great Pim Niemeijer 15-1-4 over 20 games (85% effective win rate), the first superhuman Stratego result, now published in Nature.
  • Training cost under $8,000 (16 H100s for a week plus 4 H100s for four days), against an estimated $3.0M-4.5M for DeepMind's weaker DeepNash.
  • The recipe: self-play RL, a belief network that samples plausible hidden states, and test-time search over those samples.
  • The same method set records in Barrage Stratego, Hanabi and dou dizhu, which supports the claim that it is a general pattern.
  • Training code and a CUDA simulator are on GitHub under MIT; you need a fast simulator of your own domain to reuse it outside games.

Sources: Nature: Scalable decision-making for games of imperfect information, MIT News, arXiv preprint 2511.07312, DeepNash paper (arXiv 2206.15378), AtaraxosAI/stratego on GitHub, Ataraxos game records

AIReinforcement LearningGame AIResearchImperfect InformationOpen SourceNature
CONSOLE
$