Ai2 open-sources new MoE training stack, self-tested at 2.7× throughput
Open-source training infrastructure now scales to trillion-parameter MoEs; all numbers are Ai2's own benchmarks, real-world results pending third-party use.
Ai2 released the open-source training framework Olmo-core 3 on October 1, saying it scales mixture-of-experts (MoE) training into the trillion-parameter range.
In its own test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU, versus 19,400 with the earlier implementation — about 2.7×. A separate test ran a 1.2-trillion-parameter model across 512 GPUs, but with random routing, measuring system performance only, not model quality.
The framework switches to distributed data parallelism, keeping expert weights resident on GPUs instead of repeatedly gathering them, and supports the MXFP8 low-precision format, self-tested at about 21% higher throughput than BF16. It underpins the next generation of the open Olmo models; code and a technical report are public, and real training results await third-party use.