Home/Events/Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Confirmed
Confidence
90%
Impact: 80%
Updated 37m ago

Consensus Brief

Hugging Face has released Olmo-core 3, a significant upgrade to its framework for developing large language models, featuring a redesigned open mixture-of-experts (MoE) training system. This new infrastructure is designed to scale MoE training into the trillion-parameter range while maintaining computational efficiency.

Sourced from
Primary: Hugging Face

What Changed Since Last Update

37m ago

New official source added: Hugging Face published an update on Thu, 01 Oc ("Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs").

Claim Ledger

4 claims tracked across sources

Confirmed Fact

Olmo-core 3 can scale MoE training into the trillion-parameter range.

Confirmed Fact

In a benchmark, total parameter capacity grew from 4.6B to 47B while training throughput fell by less than 5%.

Confirmed Fact

A 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack.

Confirmed Fact

Training throughput was about 21% higher with MXFP8 compared to BF16.

Role-Based Impact Analysis