Ataraxos: How a Self-Playing AI Masters Stratego, Hanabi and Dou Dizhu
Ataraxos is an artificial intelligence system designed for imperfect-information games. It combines self-play reinforcement learning, transformer networks, belief modelling and test-time search to make decisions when important information is hidden from the player.
The system was evaluated in Stratego, Hanabi and Dou Dizhu. In Stratego, Ataraxos defeated leading human players and substantially outperformed previous AI systems while using far less computational resources than DeepNash.
How Ataraxos Works
Ataraxos consists of four closely related components:
- Two interdependent self-play reinforcement-learning processes for choosing piece set-ups and moves.
- A set-up transformer that selects the initial arrangement of pieces.
- A move transformer that chooses actions during the game.
- A belief network and test-time search procedure that estimate hidden information and improve decisions before each move.
The design was developed for games such as Stratego, where players must act without seeing the identities of all opposing pieces.
Interdependent Self-Play Reinforcement Learning
Ataraxos divides Stratego into two connected learning problems. The first process learns how to select a private starting arrangement, known as a set-up. The second learns how to move pieces after the game begins.
The two processes train separately but influence each other. The set-up process provides initial boards for the move-selection process, while the move-selection process produces game outcomes used to update both policies.
This decomposition avoids forcing one end-to-end model and training pipeline to solve two substantially different tasks. Experiments found that the two phases benefit from different learning strategies:
- Set-up learning: decoder-only transformers, Monte Carlo estimates of returns and advantages, higher learning rates and higher regularization temperatures.
- Move learning: encoder-only transformers, λ-based return and advantage estimators, an annealed learning-rate schedule and lower regularization temperatures.
Transformer Networks for Stratego
Ataraxos uses two transformer networks to represent policies and value functions: a set-up network and a move network.
The set-up network uses a decoder-only architecture. This makes it possible to train on complete set-ups with a single forward and backward pass. The move network uses a key-query matrix product to parameterize the policy over legal moves, which trained faster than simpler alternatives.
Both networks use learned absolute positional embeddings. The move network was sized to balance sample efficiency with iteration speed, a trade-off that had a substantial effect on training performance.
Self-Play Training Data Generation
During self-play, Ataraxos directly samples set-ups and moves from the corresponding networks.
For moves, the system estimates expected returns and advantages using λ returns with distinct λ values. It then trains primarily on moves with large estimated advantage magnitudes. This filtering reduced the wall-clock time per reinforcement-learning iteration by approximately 2.5 times.
Filtering also increased sample efficiency based on environment queries and improved asymptotic performance. The authors note that this result was unexpected and warrants further investigation.
For set-ups, Ataraxos uses Monte Carlo returns based on the final outcomes of games played by the current policy. It does not apply advantage filtering to set-up training. The effectiveness of Monte Carlo returns for set-up advantage estimation is unusual in reinforcement learning, although similar behaviour has been observed in language-model reasoning.
Dynamically Damped Self-Play
Ataraxos trains on on-policy or nearly on-policy data by dynamically damping its learning process. This enables the system to use standard policy-optimization methods while controlling policy changes in an environment with hidden information.
Policy regularization
The set-up network uses a maximum-entropy term. The move network uses a myopic reverse Kullback–Leibler penalty toward a policy that first selects a movable piece uniformly at random and then selects one of that piece’s legal moves uniformly at random.
Ataraxos anneals the coefficients of these regularization terms according to different power laws. Regularization behaves somewhat like an energy reserve: overly cautious annealing leaves the model underdeveloped, while overly aggressive annealing can produce fast early gains followed by entropy collapse and reduced learning capacity.
Controlling update size
Ataraxos controls the size of policy updates with four complementary mechanisms:
- A reverse Kullback–Leibler penalty toward the data-collection policy.
- Importance-ratio clipping.
- Gradient-norm clipping.
- The learning rate used by Adam.
The move-network learning rate is annealed according to a power law. Scheduling the learning rate was important for fast early learning and for preventing later plateaus.
Belief Modelling for Hidden Information
To support search, Ataraxos trains a belief network on every position in trajectories generated by the final self-play policy.
Given the information available to a player, the belief network uses teacher forcing to maximize the likelihood of the true types of the opponent’s hidden pieces. It processes known information with a transformer encoder and predicts unknown pieces autoregressively in row-major order with a transformer decoder.
Dropout is applied during training to improve generalization to positions produced by opponents whose strategies differ substantially from the system’s self-play policies, including human players.
Test-Time Search with Update Equivalence
Ataraxos performs search by applying an additional damped self-play reinforcement-learning update at decision time. This approach is notable because search in games with as much hidden information as Stratego has traditionally been considered especially difficult.
- The belief network samples possible complete game states consistent with the current position.
- For each candidate move, Ataraxos runs depth-limited rollouts from those sampled states.
- The move network simulates both players during the rollouts.
- The system averages value predictions across positions reached by rollouts beginning with the candidate move.
- These values are used to update the policy with a tabular step of magnetic mirror descent.
- The move is sampled from the updated policy.
The averaged values approximate self-play action values even when the opponent does not follow Ataraxos’s policy. The belief network estimates the posterior distribution of hidden states, the rollouts use the self-play policy, and the value network predicts the value of resulting positions.
The search update uses the same two reverse Kullback–Leibler divergences used during training. It can be more aggressive than training updates because it is tabular and therefore does not alter the policy at other positions. It also benefits from more accurate advantage estimates made possible by additional computation at test time.
Ataraxos does not search when selecting its initial set-up. Set-ups are sampled directly from the set-up network.
Ataraxos Training Compute and Scale
The reinforcement-learning training run used 16 NVIDIA H100 GPUs for one week. Training the belief network used four H100 GPUs for four days.
The reinforcement-learning run included:
- 163 million completed games.
- 208 billion environment steps.
- 8.56 million move-network gradient steps.
- 99,500 set-up-network gradient steps.
Compared with earlier work that did not reach the level of top human Stratego players, Ataraxos used approximately one five-hundredth of the compute cost, one-thirtieth of the self-play games and one-hundredth of the training examples. These reductions indicate improved sample efficiency, not merely faster implementation.
Stratego Rules Explained
Board and objective
Stratego is played on a 10-by-10 board with 92 occupiable squares and two non-occupiable areas called lakes. Each player starts with 40 pieces arranged secretly on the first four rows of their side.
Players alternate turns, moving one piece at a time. When a piece moves onto a square occupied by an opponent’s piece, a battle occurs. Both pieces are revealed and at least one is removed.
A player wins by capturing the opponent’s Flag or leaving the opponent with no legal moves. A draw occurs when the player to move has no legal moves and the other player would also have no legal moves.
Piece movement and combat
Most movable pieces move one square up, down, left or right into an empty square or onto an opponent’s piece. Scouts can move any number of squares in a cardinal direction, provided they do not jump over a lake or an occupied square.
Combat is usually determined by rank. A higher-ranking piece defeats a lower-ranking piece. Pieces of equal rank are both removed.
Competitive rules
Competitive Stratego includes a two-square rule that prevents a player’s piece from crossing the same square boundary more than three consecutive times.
The continuous-chasing rule limits repeated threats during a chase. A threat moves a piece adjacent to an opponent’s piece. An evade moves a previously threatened piece away from the threatening piece. A chase is an unbroken sequence of alternating threats and evades.
A player may not make a threat that recreates a position already seen during the chase, unless the move returns the piece to the square it occupied before the previous turn of the chasing player.
Online and time-control rules
Strategus, the primary online competition website and the platform used for Ataraxos’s evaluation against Pim Niemeijer, applies a 200-move rule. A game is drawn if 200 moves occur without a battle.
The evaluation against Pim Niemeijer used the default 15+3 time control. Each player receives a 15-minute time buffer and three free seconds per move before the buffer begins to decrease.
GPU-Accelerated Stratego Simulator
To make the project feasible on modest academic computing infrastructure, the researchers implemented a GPU-accelerated Stratego simulator in CUDA C++. The simulator reduced runtime and memory usage across data collection, training and search.
The simulator was designed to:
- Avoid storing quantities that can be reconstructed when needed.
- Reach approximately 10 million board-state updates per second.
- Implement anti-chasing rules efficiently.
- Minimize dynamic memory allocation.
- Reset board states quickly for search.
- Run games independently so that terminated boards can restart without waiting for other games.
Independent resets gradually desynchronize parallel games, producing training data from different stages of play.
StrategoRolloutBuffer Design
Unlike conventional reinforcement-learning systems, Ataraxos combines the simulator and rollout buffer in one object called StrategoRolloutBuffer.
The object tracks a configurable number of parallel games and supports two essential operations:
- Historical queries, such as retrieving a legal-action mask or information-state representation from a previous position.
- Game updates through the
ApplyActionsmethod, which receives a tensor containing actions for the parallel games.
Completed games are replaced by new games. Initial boards are sampled from a configurable distribution, and terminated games can also be reset to specified non-initial positions for search.
Game histories are stored in a circular, preallocated GPU-memory buffer. Historical queries are available while the requested position remains within the buffer’s configured time range.
This design reduces memory fragmentation by reconstructing information states and legal-action masks on demand instead of copying them into a separate rollout buffer. It also supports the historical information required to enforce anti-chasing rules and to capture previous positions for search.
Implementing the two-square rule
A custom GPU state machine tracks the last four positions occupied by the last-acting piece for each player. Scouts require additional handling because of their special movement abilities.
Implementing continuous chasing
The simulator uses board history, a state-machine design and a fast differencing algorithm to determine whether a move violates the continuous-chasing rule.
Additional logic allows the rule to work when a terminated game is reset to a non-initial position. In that situation, the simulator must use the board history leading to the reset position rather than only the history stored in the rollout buffer.
Search resets
Search requires simulations to begin from a specified position. The rollout buffer supports this by resetting completed boards to custom non-initial states rather than always generating new initial boards.
Experiments and Ablation Results
The researchers conducted reinforcement-learning ablations and measured search performance across different hyperparameters. Each ablation was run from the final hyperparameter configuration without retuning the remaining settings. The results therefore measure the effect of removing or changing a component within the final design rather than the best possible version of an alternative design.
Exponential moving averages of the model parameters were also compared with current training iterates across three random seeds. The moving averages produced similar or better mean performance and lower variation between seeds.
Performance Against Stratego Bots
Ataraxos was evaluated against bots from the Stratego Evaluator benchmark:
- Asmodeus: 99 wins out of 100 games.
- Celsius: 98 wins out of 100 games.
- Celsius1.1: 97 wins out of 100 games.
- Vixen: 99 wins out of 100 games.
- Peternlewis: 98 wins out of 100 games.
Because the win rates were close to 100%, the benchmark served primarily as a sanity check.
Ataraxos’s Learned Stratego Set-Ups
Ataraxos’s set-up distribution was analyzed under the assumption that the Flag was on the left side. The probabilities on the right side are symmetric because the system randomly applies left-right orientation to sampled set-ups.
Ataraxos uses a Flag enclosed by Bombs approximately two-thirds of the time. The set-ups used in the evaluation against Pim Niemeijer were generated autonomously, with one sampled set-up per game.
Ataraxos Play Style Compared with Humans
Set-ups
Human observers found that Ataraxos frequently chooses aggressive set-ups, placing high-value pieces near the front. It also uses Bombs in the third and fourth rows and places Flags in back corners more often than humans typically do.
Its set-ups generally have less predictable structure than those used by human players.
Gameplay
Observers reported that Ataraxos makes its pieces difficult to identify from their movement patterns. It is more willing than humans to accept an opening draw when continuing would have negative expected value, as occurred in Game 11 against Pim Niemeijer.
The system tends to preserve Scouts deep into games. It uses some human-style bluffs sparingly but is willing to attempt other bluffs that people consider excessively risky.
When behind, Ataraxos fights aggressively, slowing the game and repeatedly challenging its opponent. Observers described it as particularly strong at long-term positional play, punishing mistakes, defending, playing with incomplete information, transitioning from the middle game to the endgame and converting both favourable and unfavourable positions into wins or draws.
Its gambles can appear arrogant to human players because the system does not consistently respect the opponent’s apparent strength. It can also seem unusually lucky because it frequently has the pieces it needs in the right locations and its gambles often succeed.
Ataraxos Compared with DeepNash
DeepNash is a Stratego AI developed by DeepMind. It was evaluated on Gravon in April 2022, winning 42 of 50 counted games but failing to reach the site’s top ranking.
The earlier evaluation has several limitations. By that time, many players had left Gravon, and only 25 players appeared in the final 2022 ranking. The opponents were therefore below the level of the strongest human players. The players were also unaware that an official evaluation was taking place, had no reason to search for bot-specific exploits and did not know that DeepNash used a fixed strategy that could potentially be exploited across multiple games.
At the 2023 Stratego World Championship, DeepNash was demonstrated against human players, recording 19 wins and nine losses. It lost to most of the highest-ranked players who faced it, including Pim Niemeijer.
Estimated compute cost
DeepNash was trained on 1,024 tensor-processing-unit nodes. According to the recollection of the corresponding DeepNash author, training took two to three months on TPU v3 hardware. At 2025 prices, the estimated cost would be approximately US$3 million to US$4.5 million, depending on how much of the third month was used.
Ataraxos trained its reinforcement-learning models on 16 H100 GPUs for one week and its belief model on four H100 GPUs for four days. The estimated cost was less than US$8,000 at 2025 prices.
Training sample comparison
DeepNash used approximately 5.5 billion games and between five trillion and ten trillion training examples. Ataraxos used approximately 160 million games and about 50 billion training examples.
The lower game and example counts indicate a substantial improvement in training efficiency.
Stratego Evaluation Against Top Human Players
Ataraxos was evaluated on Strategus using 5+1 time controls. Each player received a five-minute time buffer and one free second per move. Players scheduled games at their convenience and were told that Ataraxos would not adapt to their play.
Policy network results
Because belief-model training was still in progress, the first evaluation used the policy network without search. It achieved the following results:
- Against world number-four Axel Hangg: 29 wins, 17 losses and four draws.
- Against world number-three Sébastien Crot: 36 wins, 12 losses and two draws.
- Against world number-one Pim Niemeijer: 26 wins, 21 losses and three draws.
Search policy results
A second 50-game series was played against Pim Niemeijer because he had been the closest opponent in the first evaluation. The search policy had an additional disadvantage: Pim had already learned information about Ataraxos’s set-up policy and similar playing style.
Despite this, the search policy won the series with 31 wins, 14 losses and five draws. The result substantially exceeded the performance of the policy network without search.
Ataraxos in Hanabi
Ataraxos was also evaluated on two- through five-player versions of Hanabi. In the N-player version, N players used search while the remaining players did not, demonstrating how the method can scale to multi-agent search.
When all players used search, the average results over 10,000 games were:
| Players | Average score | Perfect 25-point games |
|---|---|---|
| 2 | 24.654 ± 0.007 | 77.53% ± 0.42% |
| 3 | 24.863 ± 0.005 | 89.90% ± 0.30% |
| 4 | 24.852 ± 0.005 | 88.40% ± 0.32% |
| 5 | 24.410 ± 0.009 | 58.05% ± 0.49% |
Performance improved as more players used search, with especially large gains in variants containing more players. Although the numerical increases in average score may appear small, Hanabi scores are capped at 25 and points become increasingly difficult to obtain near that ceiling. The percentage of perfect games therefore provides another measure of the substantial improvement.
Ataraxos in Dou Dizhu
Ataraxos was evaluated against PerfectDou and DouZero, two previous state-of-the-art Dou Dizhu systems. The evaluation used duplicated deals: one agent controlled the landlord while the opposing agent controlled the two independent peasants, and the roles were then swapped.
Role-averaged scores over 10,000 duplicated deals were:
- Ataraxos: 0.107 ± 0.012 against the policy network, 0.199 ± 0.015 against PerfectDou and 0.350 ± 0.016 against DouZero.
- Policy network without search: −0.107 ± 0.012 against Ataraxos, 0.152 ± 0.004 against PerfectDou and 0.286 ± 0.004 against DouZero.
- PerfectDou: −0.199 ± 0.015 against Ataraxos, −0.152 ± 0.004 against the policy network and 0.142 ± 0.004 against DouZero.
- DouZero: −0.350 ± 0.016 against Ataraxos, −0.286 ± 0.004 against the policy network and −0.142 ± 0.004 against PerfectDou.
Ataraxos achieved the strongest result against every baseline, establishing a new state of the art in the reported evaluation.
The researchers also tested different search update step sizes using performance against PerfectDou. This smaller sweep used 2,500 duplicated deals per setting.
Opportunities for Future Improvement
Learning across time
Ataraxos accesses history through features rather than learning directly across time. For the belief model, interleaving spatial attention with temporal attention or recurrent models produced substantially stronger compute-normalized performance.
The same improvement was not observed immediately for reinforcement learning because temporal architectures require additional runtime and memory and interact with advantage filtering. Such models could perform better with more compute or further runtime, memory and architectural optimization.
More powerful search
The current search procedure is ultimately limited because it imitates a single policy-update step. A more advanced search algorithm could use arbitrary additional computation to continue improving the policy.
One possible direction is to incorporate techniques from knowledge-limited subgame solving.
Ethics and Human Evaluation
The Carnegie Mellon University Institutional Review Board determined that the human-player evaluation qualified for exemption under the 2018 Common Rule, 45 CFR 46.104(d)(3)(i)(B), as a low-risk benign behavioural intervention. The study was registered as STUDY2025_00000013, with modification MOD202500000509.
All participants provided informed consent in accordance with procedures reviewed by the CMU Institutional Review Board.
Source: www.nature.com


