Artificial IntelligenceTechnical Deep Dive

Ai2 Unveils Olmo-core 3: A New Open Standard for Trillion-Parameter AI Training

Published
EElectricBuzz Editorial Team
Ai2 Unveils Olmo-core 3: A New Open Standard for Trillion-Parameter AI Training
3 min read518 wordsElectricBuzz Editorial Team

The Gist

“The Allen Institute for AI is tackling the scaling challenges of Mixture-of-Experts models with a new, highly efficient training infrastructure.”

Revolutionizing Mixture-of-Experts Training

The Allen Institute for AI (Ai2) has officially launched Olmo-core 3, a significant evolution in training infrastructure designed specifically for large-scale Mixture-of-Experts (MoE) models. As AI developers look to move beyond dense model architectures, MoEs have emerged as a leading solution for scaling parameter counts without linearly increasing computational costs. However, managing the complex routing and memory overhead associated with these models remains a hurdle. Olmo-core 3 addresses these bottlenecks, enabling developers to push model sizes into the trillion-parameter territory while maintaining high computational efficiency.

By shifting from previous fully sharded data parallelism (FSDP) implementations to a refined approach based on distributed data parallelism (DDP), the team has achieved substantial performance gains. Preliminary benchmarks conducted on NVIDIA B300 GPUs demonstrated a 2.7x increase in throughput compared to their earlier frameworks. This shift ensures that specialized experts remain resident on specific GPUs, drastically reducing the latency associated with repeatedly gathering and resharding model weights during training.

Advanced Optimization Techniques

At the heart of Olmo-core 3 is a sophisticated suite of optimization strategies designed to maximize GPU utilization. To distribute massive MoEs across hardware clusters, the framework employs three primary pillars: expert parallelism, pipeline parallelism, and distributed optimization. Expert parallelism partitions the expert pool across nodes, while pipeline parallelism breaks model layers into stages to keep memory requirements manageable. The distributed optimizer ensures that optimizer state data is spread across the GPU cluster rather than being duplicated in its entirety on every chip.

The system also introduces low-level refinements to streamline data movement and computation. Features such as GPU-resident routing metadata allow the CPU to queue tasks without waiting for round-trip data transfers, while grouped GEMM operations combine numerous smaller expert computations into larger, more efficient blocks. Additionally, support for the MXFP8 number format offers a significant efficiency boost; benchmarks indicated a 21% increase in training throughput compared to standard BF16 precision, while simultaneously reducing peak memory usage by nearly 8%. These combined techniques give researchers granular control over the complex trade-offs between speed, memory footprint, and precision.

Pushing Toward the Trillion-Parameter Horizon

The practical potential of Olmo-core 3 is evidenced by its performance at the extreme scale. During internal testing, the team successfully benchmarked the infrastructure with models containing up to 1.2 trillion parameters, maintaining an impressive throughput of 858 TFLOP/s/GPU. These tests were not just theoretical; they provided critical insights into the real-world mechanics of training large models, such as the avoidance of "token gerrymandering"—a failure mode where routing seems balanced but degrades in quality—and the nuances of overlapping communication with computation.

As an open-source project, Olmo-core 3 represents more than just a tool for the Allen Institute; it serves as a foundation for the entire research community. By providing transparent infrastructure, Ai2 aims to democratize the development of highly capable AI models. The framework is designed to be extensible, allowing developers to adapt it to diverse hardware environments and experiment with different routing or parallelism strategies. This release underscores the belief that truly open AI requires not only accessible model weights but also full transparency regarding the infrastructure and methodologies used to train them.

SPONSORED
The 5 Best Over-Ear ANC Headphones of 2026, Tested & Ranked
Editor's Pick Guide
92/100
Tech & Gadgets•12 min read

The 5 Best Over-Ear ANC Headphones of 2026, Tested & Ranked

We locked five over-ear ANC picks for 2026 — Sony WH-1000XM6, Bose QuietComfort Ultra 2, Soundcore Space One, Sennheiser Momentum 5, and Apple AirPods Max 2 — then stress-tested them on lab metrics, long-term owner truth, and live street prices.

Related Stories

Semantically matched articles, ranked by topic overlap and freshness.

Volkswagen Pivots to Wayve for Next-Gen Autonomous Driving Strategy
Artificial Intelligence

Volkswagen Pivots to Wayve for Next-Gen Autonomous Driving Strategy

In a major strategic shift, Volkswagen has reportedly selected UK-based AI firm Wayve to spearhead its autonomous driving development, edging out industry giants like Nvidia.

Photon Secures $4.5M to Funeral-March the Era of Mobile Apps
Artificial Intelligence

Photon Secures $4.5M to Funeral-March the Era of Mobile Apps

AI startup Photon has raised $4.5 million to turn its vision of agent-based messaging into a reality, betting that the future of software lies within platforms like iMessage and WhatsApp.

Legato Launches AI-Powered Smart Frames to Disrupt Hearing Tech
Artificial Intelligence

Legato Launches AI-Powered Smart Frames to Disrupt Hearing Tech

Legato is bridging the gap between stylish eyewear and medical-grade hearing assistance with its new $999 AI-integrated frames.

Shopify Unveils Canvas: Building E-Commerce Stores Through Conversational AI
Artificial Intelligence

Shopify Unveils Canvas: Building E-Commerce Stores Through Conversational AI

Shopify's new Canvas tool transforms store creation by allowing merchants to build and customize websites entirely through natural language interaction with an AI agent.

Brian Chesky: Why AI Agents Require a New Operating System
Artificial Intelligence

Brian Chesky: Why AI Agents Require a New Operating System

Airbnb CEO Brian Chesky argues that current AI interfaces fail to serve complex e-commerce, calling for a dedicated AI operating system to bridge the gap between agents and applications.

Windows 11 26H2 Update Makes Cloud-Based Settings Backup Mandatory by Default
Artificial Intelligence

Windows 11 26H2 Update Makes Cloud-Based Settings Backup Mandatory by Default

Microsoft's latest Windows 11 update shifts its cloud backup policy, leaving IT administrators scrambling to adjust settings for enterprise environments.

AWS Launches Dogwood Local Engine to Curb Rogue AI Agents
Artificial Intelligence

AWS Launches Dogwood Local Engine to Curb Rogue AI Agents

AWS has released an open-source Rust library designed to keep autonomous AI agents on a tight leash through temporal policy enforcement.

Scaling Intelligence: Hugging Face Simplifies LLM Deployment
Artificial Intelligence

Scaling Intelligence: Hugging Face Simplifies LLM Deployment

Hugging Face is streamlining the path from model selection to production-grade deployment with its latest Inference Endpoints update, featuring support for models like Salesforce’s xgen-7b-8k-base.