Article
Ai2 opens Olmo-core 3, its MoE training stack for trillion-param runs
Ai2 released Olmo-core 3, an open mixture-of-experts training framework it says reaches 2.7x the throughput of its previous implementation.
Ai2 has released Olmo-core 3, a rebuilt version of the training framework behind its Olmo models, with a new mixture-of-experts (MoE) training system aimed at models in the trillion-parameter range. The release matters less for what you can download and chat with — there is no new model here — and more for what it hands to anyone outside the big labs: a fully open stack for training the sparse architectures that now dominate frontier model design.
The company says Olmo-core 3 is one of the core systems behind the next generation of Olmo, which will use an MoE architecture, and that it is being published as part of its practice of opening the training infrastructure alongside the models, according to its announcement.
The problem Ai2 says it is solving
MoE models hold many more parameters than they use on any single input, routing each token to a handful of specialized "experts." The efficiency is real but partial. As Ai2 describes the trade-off, the full model still has to live across GPU memory and be updated during training, and sending tokens to the right experts across a cluster creates its own communication and coordination overhead. Scale the expert pool up and those costs can eat most of the advantage.
Ai2's headline claim is that Olmo-core 3 closes much of that gap. In one company-reported benchmark, the team grew the expert pool from 8 to 128 while still selecting four experts per token, holding active parameters roughly fixed at about 3.2B. Total parameter capacity went from 4.6B to 47B, and the company says training throughput fell by less than 5%.
What changed under the hood
The central architectural shift is how model weights are handled during training. Ai2's earlier MoE implementation used fully sharded data parallelism, gathering and resharding weights for each small batch. Olmo-core 3 moves to a distributed data parallelism design that keeps experts resident on GPUs and routes data to them instead, per the announcement.
The company reports that in a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack against 19,400 with the earlier implementation — roughly 2.7 times the throughput. Ai2 frames the work as bringing an integrated MoE stack to its own framework, noting that NVIDIA's Megatron-Core is the established option for this kind of training.
Three techniques split the model across hardware: expert parallelism spreads experts across GPUs, pipeline parallelism splits layers across GPU groups, and a distributed optimizer shards optimizer state rather than replicating it. On top of that sit routing optimizations — rowwise expert parallelism that writes routed data straight into expert input buffers, GPU-resident routing metadata so the CPU does not stall waiting for copies, and grouped GEMM to batch many small expert computations, all documented in the release.
Olmo-core 3 also supports MXFP8, a lower-precision number format. In a controlled benchmark on four B300 GPUs with work distributed uniformly across experts, the company says throughput was about 21% higher than its BF16 baseline while peak active memory dropped from 103 GiB to 95 GiB, with most of the gain coming from feed-forward computation and inter-expert data movement rather than attention.
The trillion-parameter numbers
Ai2 says it benchmarked a 1.2-trillion-parameter configuration with 58.36 billion active parameters per token across 512 B300 GPUs, with a highest observed throughput of 858 TFLOP/s per GPU. It also reports reaching 2.38 trillion total parameters using DeepEP v2 for inter-expert communication.
Both figures come with caveats the company states plainly: the large-scale tests used random routing to measure system performance rather than trained-model quality, and the 2.38-trillion run was a short-capacity test rather than sustained training. These are capacity demonstrations, not evidence that a model of that size trains well on this stack.
The negative results are the interesting part
Alongside the framework, Ai2's technical report documents experiments and approaches the team tested and rejected. One finding it names "token gerrymandering": a score meant to encourage balanced expert routing could improve even as the actual workload became less balanced.
Others cut against common assumptions. Lowering experts' learning rates to account for processing fewer tokens did not improve results in the model family tested. GPU calculations took different amounts of time depending on the values being processed, even at identical matrix dimensions — meaning fair performance comparisons need matching inputs, not just matching shapes. And overlapping communication with computation on separate GPU streams sometimes slowed end-to-end execution rather than speeding it up.
What is still unknown
Every performance number above is Ai2's own, measured on NVIDIA B300 hardware, and there is no independent reproduction yet. The 2.7x throughput comparison is against Ai2's previous FSDP-based implementation, not against Megatron-Core or other external baselines, so how the stack stacks up against the industry standard is an open question.
More fundamentally, the stack has not yet produced a released model. Ai2 says its next-generation Olmo will be an MoE trained on its largest dataset with its longest context window, but gives no size, date, or dataset details. Until that model ships, Olmo-core 3's value rests on its benchmarks and on whoever picks up the code and tries to train something with it.
Sources
Related
- Articledeepseek
DeepSeek Harness: an open-source, plugin-first agent harness
DeepSeek Harness is in public preview: an open-source agent harness where tools, the interface and even the agent loop are swappable plugins.
3 min read