Aleph Alpha Releases Kolibri: A 78.1B Open-Weight English-German MoE Model With Only 3.46B Active Parameters
78.1B parameters, 3.46B active. Aleph Alpha's Kolibri runs on a single GPU—and handles English and German at 1M-token context.

Why it matters
Aleph Alpha's Kolibri demonstrates efficient MoE scaling for multilingual reasoning. Open-weight release with FP8 quantization and single-GPU deployability lowers the barrier for enterprise practitioners to experiment with bilingual, long-context models without large distributed infrastructure.
The key facts
6 to know78.1B total parameters; 3.46B active per token (MoE design)
1M-token context window
English-German bilingual capability
Apache 2.0 license; FP8 weights
Deployable on single NVIDIA B200 or H200 GPU
Per-request reasoning effort configurable
The story so far
Earlier coverage of this storyline
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Aleph Alpha has released Kolibri, a 78.1B-parameter English-German Mixture-of-Experts model that activates only 3.46B parameters per token. It has a 1M-token context and per-request reasoning effort, and its Apache 2.0 FP8 weights run on a single B200 or H200.