The Allen Institute for AI has released Olmo-core 3, a significant upgrade to its framework for developing large language models. The update features a redesigned training system specifically optimized for Mixture-of-Experts (MoE) architectures, which allow models to contain more parameters while maintaining computational efficiency by only activating a subset of them for each input.
Model ReleasesAllen Institute for AIOlmo-core 3
Allen AI Releases Olmo-core 3 with Scalable MoE Training Infrastructure
Olmo-core 3 is designed to scale MoE training into the trillion-parameter range. In benchmarks, the framework successfully increased the expert pool from 8 to 128 while keeping the number of active parameters per token fixed at approximately 3.2B, while maintaining training throughput.
The new infrastructure switches from a fully sharded data parallelism (FSDP) approach to a system based on distributed data parallelism (DDP). This change keeps experts resident on GPUs and routes data to them, avoiding repeated weight gathering. Preliminary tests on eight NVIDIA B300 GPUs showed that a 47-billion-parameter MoE achieved approximately 2.7× the throughput compared to the previous implementation.
The framework also incorporates several optimizations, including rowwise expert parallelism, GPU-resident routing, and grouped GEMM to improve computation efficiency. Additionally, Olmo-core 3 supports the MXFP8 lower-precision number format, which can reduce computation costs and data movement between GPUs. Tests indicated that enabling MXFP8 could increase training throughput by approximately 21% compared to the BF16 baseline on NVIDIA B300 GPUs.
Olmo-core 3 serves as the foundation for the next generation of Olmo models, which are expected to utilize MoE architectures to achieve higher capabilities. The framework is released as an open-source tool for researchers and developers to train their own MoE models and experiment with various routing and parallelism configurations.
Sources
- Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs (Hugging Face Blog, 2026-10-01)
- Code