Why training cost hides in the plumbing
The arithmetic of AI training looks simple from the outside: parameters times data times hardware. In practice, a large share of the cost hides in logistics. An MoE model can contain far more learned components than a dense model of equal compute, because each input uses only a handful of experts. But the whole model still has to live across GPU memory and be updated during training, and directing each token to the right experts across a cluster creates communication and coordination overhead. As Ai2's announcement puts the problem, "As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input."
That overhead is one reason advanced model development has concentrated in a handful of well-funded labs. A framework that reduces it does not make models smarter by itself, but it changes who can afford to build one.
What Olmo-core 3 actually changes
The headline engineering change is a redesign of how the framework handles expert parallelism. Ai2's earlier MoE implementation used fully sharded data parallelism (FSDP), which repeatedly gathers and reshards model weights for each small batch of training data. Olmo-core 3 switches to distributed data parallelism (DDP), keeping experts resident on their GPUs and routing data to them, avoiding that repeated weight gathering.
The company's own preliminary test illustrates the difference: on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation, about 2.7 times the throughput. Three distribution techniques underpin the scaling: expert parallelism, which spreads experts across GPUs; pipeline parallelism, which splits model layers across GPU groups; and a distributed optimizer, which spreads optimizer state instead of duplicating it on every machine. Routing optimizations, including placing routed data directly into expert input buffers, keeping routing metadata on the GPUs, and grouping many small expert computations into larger ones (grouped GEMM), reduce the tax of deciding where each token goes.
The framework also supports MXFP8, a lower-precision number format. In Ai2's controlled benchmark on four B300 GPUs, enabling MXFP8 where it helped most produced about 21 percent higher training throughput than BF16 while peak active memory fell from 103 GiB to 95 GiB.
The scaling numbers, and what they do not show
The scaling evidence is Ai2's own and should be read as system benchmarks, not trained-model results. In one configuration, the team increased the expert pool from 8 to 128 experts while still selecting only four per token, holding active parameters roughly fixed at about 3.2 billion per token. Total parameter capacity grew from 4.6 billion to 47 billion while training throughput fell by less than 5 percent. That is the core pitch: many more parameters without proportionally more per-token compute or much slower training.
At the top end, the team benchmarked a 1.2-trillion-total-parameter model with 58.36 billion parameters active per token across 512 GPUs, observing up to 858 TFLOP/s per GPU. Critically, these tests used random routing to measure system performance rather than the quality of a trained model. An experiment with DeepEP v2, an alternative inter-expert communication method, reached a configuration with 2.38 trillion total parameters, but Ai2 describes that as a short-capacity test demonstrating reachable scale, not sustained training performance. Readers should treat the trillion-parameter figures as infrastructure capability demonstrations, not as evidence that a working model of that size has been trained.
The measurement pitfalls worth more than the benchmarks
Arguably the most valuable part of the release is not the speedups but the frank documentation of measurement traps. The technical report describes a failure Ai2 named "token gerrymandering": a score intended to encourage balanced routing could improve even as the actual workload became less balanced. In other words, a metric that looks healthy on paper can coexist with a system that is getting worse, a lesson for anyone judging AI performance claims, including claims from vendors.
Other reported findings: lowering experts' learning rates because they see fewer tokens did not improve results in the model family tested; identical GPU calculations took different amounts of time depending on the values processed, so performance comparisons need matching input values, not just matching shapes; and overlapping communication and computation on separate GPU streams, a common optimization, sometimes slowed end-to-end execution. Each of these is the kind of unglamorous negative result that rarely makes an announcement and that open infrastructure makes shareable.
What it means for who can build AI
Ai2 states that its next-generation Olmo will use an MoE architecture, aiming for its most capable Olmo yet with its largest dataset and longest context window. The lineage runs from OlmoE, with 64 routed experts, through the dense Olmo 3, to this MoE-native training stack. For researchers and smaller labs, the significance is the stated openness: the framework, code and report are released so others can train their own MoEs, adapt the stack to different hardware and experiment with routing and parallelism.
The unverified part is how well this transfers beyond Ai2's environment. All benchmarks here are vendor-run on NVIDIA B300 hardware, and a 2.7-times stack improvement is not a model-quality claim. Independent reproduction on other clusters will tell whether the gains hold. What can be said with confidence today is that a documented, open path now exists for building much larger sparse models without proportionally larger clusters, and that the tools behind an open model are, for the first time in this lineage, as open as the model itself.
