Revolutionizing Mixture-of-Experts Training
The Allen Institute for AI (Ai2) has officially launched Olmo-core 3, a significant evolution in training infrastructure designed specifically for large-scale Mixture-of-Experts (MoE) models. As AI developers look to move beyond dense model architectures, MoEs have emerged as a leading solution for scaling parameter counts without linearly increasing computational costs. However, managing the complex routing and memory overhead associated with these models remains a hurdle. Olmo-core 3 addresses these bottlenecks, enabling developers to push model sizes into the trillion-parameter territory while maintaining high computational efficiency.
By shifting from previous fully sharded data parallelism (FSDP) implementations to a refined approach based on distributed data parallelism (DDP), the team has achieved substantial performance gains. Preliminary benchmarks conducted on NVIDIA B300 GPUs demonstrated a 2.7x increase in throughput compared to their earlier frameworks. This shift ensures that specialized experts remain resident on specific GPUs, drastically reducing the latency associated with repeatedly gathering and resharding model weights during training.
Advanced Optimization Techniques
At the heart of Olmo-core 3 is a sophisticated suite of optimization strategies designed to maximize GPU utilization. To distribute massive MoEs across hardware clusters, the framework employs three primary pillars: expert parallelism, pipeline parallelism, and distributed optimization. Expert parallelism partitions the expert pool across nodes, while pipeline parallelism breaks model layers into stages to keep memory requirements manageable. The distributed optimizer ensures that optimizer state data is spread across the GPU cluster rather than being duplicated in its entirety on every chip.
The system also introduces low-level refinements to streamline data movement and computation. Features such as GPU-resident routing metadata allow the CPU to queue tasks without waiting for round-trip data transfers, while grouped GEMM operations combine numerous smaller expert computations into larger, more efficient blocks. Additionally, support for the MXFP8 number format offers a significant efficiency boost; benchmarks indicated a 21% increase in training throughput compared to standard BF16 precision, while simultaneously reducing peak memory usage by nearly 8%. These combined techniques give researchers granular control over the complex trade-offs between speed, memory footprint, and precision.
Pushing Toward the Trillion-Parameter Horizon
The practical potential of Olmo-core 3 is evidenced by its performance at the extreme scale. During internal testing, the team successfully benchmarked the infrastructure with models containing up to 1.2 trillion parameters, maintaining an impressive throughput of 858 TFLOP/s/GPU. These tests were not just theoretical; they provided critical insights into the real-world mechanics of training large models, such as the avoidance of "token gerrymandering"—a failure mode where routing seems balanced but degrades in quality—and the nuances of overlapping communication with computation.
As an open-source project, Olmo-core 3 represents more than just a tool for the Allen Institute; it serves as a foundation for the entire research community. By providing transparent infrastructure, Ai2 aims to democratize the development of highly capable AI models. The framework is designed to be extensible, allowing developers to adapt it to diverse hardware environments and experiment with different routing or parallelism strategies. This release underscores the belief that truly open AI requires not only accessible model weights but also full transparency regarding the infrastructure and methodologies used to train them.









