Actualités1 min de lecture·
Ai2 open-sources MoE training stack, self-tested at 2.7× throughput
Training very large mixture-of-experts models gets cheaper for open labs, but every number is Ai2's own benchmark.
ImportanceMatérielPreuvesE2 non réplicableTraitementRapide
Ai2 released the open-source training framework Olmo-core 3 on October 1, with self-tested throughput 2.7× its previous implementation.
In the official benchmark, a 47-billion-parameter mixture-of-experts model processed 52,000 tokens per second per GPU on eight NVIDIA B300s, versus 19,400 on the earlier stack. The trillion-parameter tests used random routing and measure system performance, not training quality.
Code and a technical report are public. Academic and smaller labs gain another option for training very large MoEs, and the numbers come from Ai2's own tests.