Aleph Alpha · 2026-10-03 · major
Kolibri — Aleph Alpha's open-weight German-English model
Kolibri is an Apache-2.0 mixture-of-experts model from Aleph Alpha with 78B total and 3.46B active parameters. It is built for German and English, scores 96.9 on AIME 2025 and handles up to 1M tokens of context.

A 78B open-weight MoE model trained in Germany and Finland, tuned for German and English reasoning.
Key specs
| Active params | 3.46B |
|---|---|
| Aime 2025 | 96.9 |
Quick facts
| Maker | Aleph Alpha |
|---|---|
| Parameters | 78B total, 3.46B active per token |
| Architecture | MoE, 50 layers, 384 experts (6 active) |
| Context window | Up to 1,048,576 tokens (trained to 256K) |
| Languages | German and English |
| License | Apache-2.0 |
| Availability | Open weights on Hugging Face (FP8 and BF16) |
Benchmarks
What is it?
Kolibri is Aleph Alpha's new open-weight reasoning model for German and English, released on October 3, 2026 under Apache 2.0. It supports tool calling, an explicit reasoning mode and very long inputs, and comes as FP8 weights plus a BF16 variant.
How does it work?
The model is a mixture-of-experts transformer: each of its 50 layers holds 384 experts, and only 6 of them run for each token. That keeps the per-token compute at about 3.46B parameters out of 78B in total. Aleph Alpha trained the model from scratch, with 21.3% of the pre-training data in German.
Why does it matter?
European teams get a permissively licensed reasoning model with German as a core training language that they can host themselves. In Aleph Alpha's own tests, Kolibri beats Qwen3.6-35B, Nemotron 3 Super and Mistral Small 4 on AIME 2025, GPQA Diamond and LiveCodeBench v6.
Who is it for?
European enterprises, public sector teams and German-language developers
Frequently asked questions
- Is Kolibri open source and free for commercial use?
- Kolibri's weights and configuration files are released by Aleph Alpha under the Apache 2.0 license. Apache 2.0 is a permissive license that allows commercial use, changes and redistribution. The full weights can be downloaded from the Aleph-Alpha/Kolibri-1 page on Hugging Face, with an FP8 version and a separate BF16 variant.
- How does Kolibri compare to Qwen3.6-35B and Mistral Small 4?
- Aleph Alpha's own table puts Kolibri at 96.9 on AIME 2025, against 84.6 for Qwen3.6-35B and 79.8 for Mistral Small 4. On LiveCodeBench v6 Kolibri scores 85.9 versus 82.5 and 71.2. On HumanEval+ it is level with the others at 92.7, and Nemotron 3 Super leads that test with 94.7.
- How long a context can Kolibri handle?
- Kolibri was trained on contexts of up to 256K tokens and supports up to 1,048,576 tokens, according to Aleph Alpha's model card. The card recommends 262,144 tokens when efficiency matters. Long-context processing is one of the three areas the Kolibri model card lists, next to the explicit reasoning mode and tool calling.
- Why does Aleph Alpha call Kolibri a sovereign model?
- Aleph Alpha says Kolibri was trained from scratch on infrastructure in Germany and Finland, under European and German law. About 21.3% of its pre-training tokens are German. Aleph Alpha describes Kolibri as a specialized sovereign model for business-critical environments, and the open Apache 2.0 weights let organisations run it on their own hardware.
Try it
pip install 'aleph-alpha-inference>=1' && vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice