huggingface.co web signal

Aleph Alpha releases Kolibri-1, a 78B German-English MoE

TL;DR

  • Kolibri-1 shipped October 3, 2026 as a 78B-parameter Mixture-of-Experts with 3.46B active per token, weights released under Apache 2.0.
  • The 20-trillion-token pre-training mix is 43.4% English and 23.4% German; the card frames this as depth over multilingual breadth.
  • Overall English score is 75.5 vs 80.2 for the card's unnamed best-dense baseline; GPQA Diamond lands 84.3 and SWE-Bench Verified 66.4.

Aleph Alpha released Kolibri-1 on Hugging Face on October 3, 2026, a 78-billion-parameter Mixture-of-Experts model with weights published under Apache 2.0. Only 3.46 billion parameters are active per token.

The pitch is bilingual depth rather than global coverage. The model card states the choice directly: "Depth over breadth supporting two languages excellently rather than many languages adequately." The training mix tracks that logic. English web and documents make up 43.4% of the 20-trillion-token pre-training corpus, German another 23.4%, and the tokenizer, UniBPE, was built around German morphology.

Pre-training ran on 768 NVIDIA B200 GPUs for 21 days: 392,000 GPU-hours, 6.4e23 FLOPS, and an estimated 9.5x10^2 MWh of energy across pre-training, mid-training and the long-context extension.

The benchmarks are published honestly. Kolibri scores 84.3 on GPQA Diamond in English and 81.3 in German, 96.9 on AIME 2025, 66.4 on SWE-Bench Verified. The card's own "Best Dense Model" column outscores Kolibri on every category: overall 80.2 vs 75.5 in English, 79.9 vs 70.8 in German. The card never names which dense model it is benchmarking against. On RULER at a 1-million-token context, Kolibri reaches 63.2%; the best cited baseline sits at 73.1% but only at 512k.

Aleph Alpha frames Kolibri as sovereign infrastructure. The company is a signatory of the EU GPAI Code of Practice, and the card lists autonomous operation without human review as explicitly outside intended use. It also concedes a specific inheritance problem: "training data included material from Chinese language models with known political bias; this was actively reduced through data filtering and dedicated alignment training."

Three analysts in our tracker shared the Hugging Face page on launch day.

Shared on Bluesky by 3 AI experts