October 1, 2026
MoE training gets faster: Olmo-core 3 delivered a 2.7x speedup in an AllenAI test
AllenAI released Olmo-core 3 on Oct 1: infrastructure for training MoE models that routes data to experts on GPUs.

In an initial AllenAI test, a 47 billion parameter model sped up from 19,400 to 52,000 tokens per second per GPU. The measurement used 8 NVIDIA B300 GPUs.
Olmo-core is for training models, rather than generating code in a chat. In MoE, each token is processed by part of the model: selected experts, rather than all parameters at once.
Experts stay in place. The previous FSDP setup repeatedly gathered expert weights for computation. In Olmo-core 3, AllenAI switched to DDP: experts stay on GPUs, and the system routes data to them.
Increasing the number of experts from 8 to 128 expanded model capacity from 4.6 to 47 billion parameters. Speed dropped by less than 5%. Each token used 4 experts with a total of 3.2 billion parameters.
Results from other tests
In a controlled test on 4 NVIDIA B300 GPUs with a balanced workload, the MXFP8 format increased speed by about 21% compared with BF16. Peak memory usage dropped from 103 to 95 GiB.
AllenAI tested a 1.2 trillion parameter model on 512 NVIDIA B300 GPUs. Each token used 58.36 billion active parameters. With random expert selection, the team reached a peak of 858 trillion operations per second per GPU.
How to install Olmo-core
Installation with pip. First install PyTorch for your hardware, then run `pip install ai2-olmo-core`. For installation from the repository, the README provides this sequence:
```sh git clone https://github.com/allenai/Olmo-core.git cd Olmo-core pip install -e '.[all]' ```
Dropless MoE, where the system does not discard tokens when an expert is overloaded, requires `grouped_gemm`. The README warns that until the fix in PR #21 is released after v0.1.6, the dependency may need to be built from source.
Prebuilt Docker images contain the core and optional dependencies, but Olmo-core must be installed separately in them. AllenAI tests these images on H100 clusters. Version v3.0.0 is dated Sep 30, 2026, and the announcement was published on Oct 1.
Original source: [AllenAI — how Olmo-core 3 works and test results, Oct 1, 2026](https://huggingface.co/blog/allenai/olmocore3).
With DeepEP v2, AllenAI fit a 2.38 trillion parameter model in a short capacity test, without a full training run.
