October 1, 2026
Ataraxos beats Stratego's most decorated player: 15 wins against Pim Niemeijer
A separate neural network estimates the opponent's hidden pieces, while Ataraxos evaluates possible lines of play before making a move.

Ataraxos won 15 games against Pim Niemeijer, lost one and drew four. The result was published on Sep 30, 2026.
In Stratego, players cannot see the types of their opponent's pieces until they clash. Ataraxos first learns through self-play, then trains a belief network to estimate hidden pieces from the available information. Before making a move, the system samples likely piece arrangements and evaluates how play might continue.
Less compute. Compared with DeepNash, the authors used roughly 500 times less compute and 30 times fewer completed training games. Ataraxos trained on 163 million games, with an estimated training cost of less than $8,000 at 2025 prices.
The main training took a week on 16 NVIDIA H100 GPUs. The network for estimating hidden pieces trained for another four days on four H100 GPUs. During evaluation, Ataraxos took an average of 1.26 seconds per move on a single H100.
Code to run. The authors released the implementation and a training demo. Linux and a GPU supporting CUDA 11.8 are required. From the repository directory, install it as follows:
```bash micromamba env create -f environment.yml micromamba activate ataraxos make pip install -e . ```
Run the demo with the command below. On an RTX A6000, one iteration typically takes 10โ20 seconds.
```bash python scripts/train/rl_main.py --num_envs 128 --rl.dtype float32 --rl.torch_compile false ```
Primary sources: [Ataraxos study in Nature, Sep 30, 2026](https://www.nature.com/articles/s41586-026-11036-y), [code and setup instructions](https://github.com/AtaraxosAI/stratego).
The same approach set a new best result in Hanabi for 2โ5 players and beat the PerfectDou and DouZero bots in dou dizhu.
