Ai2 releases Olmo-core 3 to make developing large mixture-of-experts LLMs more efficient
Allen Institute for AI releases Olmo-core 3: a training framework that scales mixture-of-experts models to a trillion parameters without proportional cost increases.

Why it matters
Olmo-core 3 addresses a concrete MoE training bottleneck — enabling trillion-parameter models while preserving computational efficiency. For practitioners building or evaluating large sparse models, this is a capability and cost development; for enthusiasts, it represents progress in making frontier-scale training accessible beyond well-capitalized labs.
The key facts
6 to knowAllen Institute for AI (Ai2) released Olmo-core 3 on October 2, 2026
Framework targets mixture-of-experts (MoE) LLM training at trillion-parameter scale
Claims to preserve computational efficiency while scaling MoE training
MoE architecture allows selective activation of model parameters (routing-based sparsity)
Source does not disclose benchmark comparisons, measured cost reductions, or independent validation
Open-source status and availability terms not stated in excerpt
Go to the source
SiliconAnglesiliconangle.com
Publisher excerpt: Seattle-based artificial intelligence research firm Allen Institute for AI announced a development framework for large language models Thursday that significantly improves how mixture-of-experts large language models are trained. The new framework, Olmo-core 3, allows MoE training to reach the…