Allen Institute Releases Olmo-Core 3 for Trillion-Parameter MoE Training

The Allen Institute for Artificial Intelligence (Ai2) has released Olmo-core 3, an open-source training infrastructure framework designed to scale mixture-of-experts (MoE) language models into the trillion-parameter range while maintaining computational efficiency. The framework serves as the foundational training stack for the next generation of Ai2's Olmo models, continuing the organization's commitment to open-source model development and training infrastructure transparency.
According to Allen Institute for AI, for Ai2 Olmo development history, Ai2's training infrastructure evolved from OlmoE's 64-expert MoE architecture through the dense Olmo 3 model, leading to the redesigned Olmo-core 3 training stack. According to Allen Institute for AI, olmo-core 3 brings an integrated, fully open-source MoE training stack to the Olmo framework, offering an open alternative to NVIDIA's established Megatron-Core system.
Training large language models demands substantial computational power, driving up energy costs and restricting advanced research. While MoE architectures offer efficiency by activating only specialized sub-components or parameters for each input token, scaling them across GPU clusters introduces communication and coordination bottlenecks. Olmo-core 3 is engineered to mitigate these synchronization overheads, enabling researchers to expand model capacity without proportional losses in training throughput.
Architectural Redesign and Parallelism Strategies
Transitioning from earlier implementations that relied on fully sharded data parallelism (FSDP), Olmo-core 3 adopts a distributed data parallelism (DDP) system. This approach maintains expert weights resident directly on GPUs and routes relevant input data to them, eliminating repeated weight gathering across training batches. In preliminary benchmarks conducted on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE model achieved 52,000 tokens per second per GPU, representing approximately 2.7 times the throughput of earlier FSDP iterations.
The framework combines three core hardware distribution techniques: expert parallelism to distribute expert pools across GPUs, pipeline parallelism to split model layers across GPU groups, and a distributed optimizer to partition training state memory. Additionally, rowwise expert parallelism places routed data directly into expert input buffers, while GPU-resident routing and grouped GEMM computations optimize hardware execution efficiency.
Precision Formats and Trillion-Parameter Scalability
Olmo-core 3 incorporates support for MXFP8, a lower-precision number format that reduces memory usage and data movement across GPUs. Controlled benchmarks on four NVIDIA B300 GPUs demonstrated a 21 percent increase in end-to-end training throughput when using MXFP8 compared to BF16 baselines, alongside a reduction in peak active memory from 103 GiB to 95 GiB.
According to Allen Institute for AI, for Ai2 MoE research team, Ai2 identified the 'token gerrymandering' phenomenon where routing scores masked unbalanced expert workloads, and discovered that overlapping communication and computation on separate GPU streams sometimes reduced throughput.
Ai2 has successfully benchmarked the infrastructure across expanded configurations, including a 1.2-trillion-parameter model featuring 58.36 billion active parameters per token across 512 NVIDIA B300 GPUs, achieving peak observed throughput of 858 TFLOP/s/GPU. Experimental testing with DeepEP v2 further demonstrated short-capacity scaling up to 2.38 trillion parameters. The technical release also documents empirical findings on routing dynamics, including the identification of token gerrymandering and communication-computation overlap trade-offs.
Sources & Citations
- Hugging Face Blog PostPrimary / official
- Allen Institute for AI
