AI2 Ships Olmo-core 3, Claims 2.7x MoE Throughput vs Megatron
TL;DR
- AI2's Olmo-core 3 processed 52,000 tokens per second on Nvidia B3000 GPUs for a 47 billion-parameter mixture-of-experts model, in AI2's own benchmarks.
- On the same workload, Nvidia's Megatron-core training stack topped out at about 19,400 tokens per second, roughly 2.7x slower.
- The framework scales the expert pool from eight to 128 while selecting four experts per token, and AI2 says the design reaches more than one trillion parameters.
The Allen Institute for AI released Olmo-core 3 on Thursday, a redesigned open framework for training mixture-of-experts language models that the Seattle lab says reaches trillion-parameter scale without the usual efficiency penalty, SiliconAngle reports.
In AI2's own benchmarks, the framework "processed 52,000 tokens per second on Nvidia B3000 GPUs for a 47 billion-parameter model." Nvidia's own Megatron-core training architecture, which the article calls "an established option for training large MoEs," topped out at about 19,400 tokens per second on the same workload, roughly 2.7 times slower. The design leans on expert parallelism to spread experts across multiple GPUs, a distributed optimizer that splits optimizer state across cards, and support for the MXFP8 number format, which the write-up describes as a way to "reduce the computation and amount of data moved between GPUs."
Olmo-core 3 "was built to bridge the gap between dense models and MoE models, allowing the expert pool to grow from eight to 128 while still selecting only four experts per token," and AI2 says the same infrastructure can scale past one trillion parameters. The code is on GitHub. The benchmark comparison is AI2's own, with no independent validation of the 2.7x figure yet, and this release lands amid a steady run of open-training infrastructure work on our open source tracker.
Originally reported by siliconangle.com
Read the original article →Original headline: AI2 Ships Olmo-core 3, Open MoE Training Stack Hits 2.7x Throughput Over FSDP at Trillion-Parameter Scale