AllenAI releases Olmo-core 3 for open mixture-of-experts training
AllenAI released Olmo-core 3, an open training stack that distributes mixture-of-experts routing and optimizer work and supports selective MXFP8 computation.
AllenAI released Olmo-core 3, an open upgrade to its language-model training framework for large mixture-of-experts (MoE) models. The system replaces the project’s earlier Fully Sharded Data Parallel path for MoE training with a Distributed Data Parallel design that keeps experts resident on GPUs and routes tokens to the experts selected for each input.
The redesign targets the repeated weight movement AllenAI identified in its previous implementation. It also gives researchers an open stack that combines expert parallelism, pipeline parallelism and sharded optimizer state for large MoE training runs.
In an MoE model, only a subset of specialized parameter groups—experts—processes each token. Olmo-core 3 uses expert parallelism to spread those experts across devices and pipeline parallelism to divide model layers among GPU groups. A distributed optimizer shards optimizer state.
With rowwise expert parallelism, Olmo-core 3 writes routed token data into expert input buffers while keeping routing metadata on the GPU. Grouped general matrix multiplication, or grouped GEMM, combines many small expert computations into larger operations. Under the earlier FSDP approach, expert weights were repeatedly gathered and resharded for each microbatch; the new DDP-based path leaves those weights in place and moves token data instead.
AllenAI said an expert-scaling benchmark increased the pool from eight experts to 128 while continuing to select four experts per token. Active parameters remained near 3.2 billion as total model capacity rose from 4.6 billion to 47 billion parameters, with a throughput reduction of less than 5%, according to the organization.
In a separate preliminary test on eight NVIDIA B300 GPUs, AllenAI reported 52,000 tokens per second per GPU for a 47-billion-parameter MoE, compared with 19,400 for its earlier implementation. The organization described that as roughly 2.7 times the throughput. That training test is not directly comparable with reviewed NVIDIA MLPerf inference results, which cover a different workload.
Olmo-core 3 also supports MXFP8 for selected operations. The lower-precision format can reduce computation and communication costs, although AllenAI said conversion overhead can erase the gain in some cases. In a controlled four-B300 benchmark with uniform expert workloads, the organization said selective MXFP8 increased end-to-end training throughput by about 21% over BF16 and reduced peak active memory from 103 GiB to 95 GiB. The public repository separately documents float8 support through torchao.
For its largest reported systems benchmark, AllenAI tested a 1.2-trillion-parameter configuration with 58.36 billion active parameters per token across 512 B300 GPUs and measured a peak 858 TFLOP/s per GPU. The test used random routing to measure systems performance, not model quality. AllenAI also said a short DeepEP v2 capacity experiment reached 2.38 trillion total parameters, but did not describe it as a sustained training run. None of those performance results was independently reproduced in the reviewed sources.
AllenAI said its next-generation Olmo model will use an MoE architecture, a stated plan rather than a released model. The Olmo-core code is publicly available under the Apache 2.0 license, with installation documented from source or through the ai2-olmo-core Python package.
More news

Databricks adds branch-based restores to Lakebase Postgres

AWS adds hierarchy filtering to Amazon Quick Sight dashboards

ServiceNow CoreAI introduces AutoSynthData for enterprise-agent training
