On 2026-10-01 Ai2 released Olmo-core 3, an open MoE training stack (GitHub, PyPI ai2-olmo-core). It moves from FSDP to DDP + expert parallelism; on 8×B300 a 47B MoE hits ~2.7× throughput (52k vs 19.4k tok/s/GPU). Systems benches include 1.2T on 512 GPUs at up to 858 useful-model TFLOP/s/GPU and a DeepEP v2 capacity probe at 2.38T. Report: olmocore3.

Key Takeaways

  • ✓Announced 2026-10-01: allenai.org/blog/olmocore3 + HF mirror + github.com/allenai/OLMo-core
  • ✓Stack: EP/PP/distributed optimizer; NVSHMEM rowwise EP + GPU-resident routing + grouped GEMM; optional DeepEP v2
  • ✓Throughput: expert pool 8→128 at ~3.2B active, <5% drop; 47B MoE on 8×B300 52k vs 19.4k tok/s/GPU (~2.7×)
  • ✓Scale: 1.2T / 58.36B active / 512 GPUs up to 858 TFLOP/s/GPU (random-routing systems ref); DeepEP v2 probe 2.38T
  • ✓MXFP8 ~+21% vs BF16, peak active mem 103→95 GiB; topology-agnostic checkpoints; next Olmo will be MoE
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points

Frontier MoEs grow total capacity far faster than per-token active compute, but dense-era full-reshard FSDP still gathers weights by total expert capacity each microbatch—Ai2’s “capacity tax.” Labs need an open stack that still works near trillion scale. On 2026-10-01 Ai2 shipped Olmo-core 3 (HF mirror) as infrastructure for the next MoE Olmo.

Architecture Highlights & Internals

The new path centers on DDP: keep experts resident and route token rows. Compose EP, PP, and a distributed optimizer. Systems pieces: NVSHMEM rowwise EP, GPU-resident routing, device-scheduled grouped GEMM, MXFP8, topology-agnostic FP32 checkpoints, optional DeepEP v2. Code: allenai/OLMo-core / PyPI ai2-olmo-core. Report: olmocore3.

Authoritative Benchmarks & Measured Scores

Figures are Ai2 systems numbers under random routing—not model quality. Capacity sweep: top-4, ~3.2B active, experts 8→128, throughput drop <5%. 8×B300 47B MoE: ~52k vs ~19.4k tok/s/GPU (~2.7×). 1.2T / 58.36B active / 512 GPUs up to 858 TFLOP/s/GPU; DeepEP v2 probe 2.38T. MXFP8: ~+21% vs BF16, peak active mem 103→95 GiB. Report also documents token gerrymandering and failed overlap tricks—don’t treat as a Megatron replacement claim.

Developer Hands-on Guide

Read blog + report §14 conditions; install ai2-olmo-core or clone GitHub; validate EP/PP/MXFP8/recompute on small runs before scaling; compare with Megatron-Core on your topology; this is pretrain infra, not an inference agent model.