Overview
Kolibri is a Mixture-of-Experts language model from the German company Aleph Alpha, released on October 3, 2026 (the Day of German Reunification) with full weights on Hugging Face under the Apache 2.0 license. It has 78.1B total parameters, of which 3.46B are active per token, and it is built for German and English. The model supports an explicit reasoning mode with four effort levels (none, low, medium, high) and tool calling, and Aleph Alpha positions it for regulated work in areas such as public administration, industrials and aerospace, run on-premise.
Kolibri was pre-trained on sequences of 16,384 tokens, mid-trained at 65,536 and trained at 262,144 tokens in a final long-context phase, which the model card calls its native context length. Because positional encoding is applied only in the sliding-window layers, Aleph Alpha says the context can be extended without position scaling, and it has validated quality and serving up to 1,048,576 tokens; it recommends at most 262,144 tokens for latency-sensitive deployments and complex tasks. Of its 50 layers, 40 use a 512-token sliding window and 10 attend across the full context.
Training ran on 768 NVIDIA B200 GPUs in three stages: 20T tokens of pre-training over 21 days, 3.44T tokens of mid-training and about 200B tokens of long-context adaptation. German made up 21.3% of pre-training tokens (about 4.3T), against roughly 62% English and 14% code, and Aleph Alpha built a bilingual English-German tokenizer (128,000 vocabulary, trained with a method it calls UniBPE) so German text needs fewer tokens. Post-training combined supervised fine-tuning on a 268B-token mix with reinforcement learning on more than 1.2 million curated tasks, plus abstention training with Aleph Alpha's Merlin-Arthur procedure so the model says it does not know when the context lacks the answer.
In the model card's post-training table, Kolibri scores 84.3 on GPQA Diamond, 96.9 on AIME 2025, 85.9 on LiveCodeBench v6, 66.4 on SWE-Bench Verified and 80.0 on MMLU-Pro (CoT), with German variants alongside (81.3 on German GPQA Diamond, 87.5 on German AIME 2025). In the launch post's comparison it leads Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B on AIME 2025, GPQA Diamond and τ³-bench banking, while Qwen3.6-35B-A3B leads on BFCL v4, τ²-bench telecom and LongBench Pro. Kolibri succeeds Kolibri Origin, a 30.6B-total internal model that finished pre-training on June 11, 2026 and was not publicly released.
Serving requires Aleph Alpha's aleph-alpha-inference package, a vLLM plugin published on GitHub and as a container image. The released checkpoint uses FP8 weights with a memory footprint of about 78 GB; the model card lists 2× A100 80 GB, 2× H100 SXM5, or a single H200, B200 or B300 as the minimum hardware, and the vLLM server exposes an OpenAI-compatible API.
| Released | 2026-10-03 |
|---|---|
| License | Apache-2.0 |
| Weights | Open weights |
| Parameters | 78.1B total · 3.46B active |
| Context | 1,048,576 tokens (validated); 262,144 native and recommended |
| Architecture | Mixture-of-Experts transformer: 50 MoE layers with 384 experts per layer (1 shared + 6 routed active), sliding-window attention (512 tokens) with full attention every 5th layer, trained with Muon and exact quantile balancing |
| Knowledge cutoff | June 18, 2026 (English and German) |
| Modalities | Text |
Benchmarks
Kolibri against Kolibri Origin and larger open MoE models, as published in Aleph Alpha's launch post (October 3, 2026). Shared 0–100 scale, higher is better; the AA-Omniscience Index runs from −100 to 100.
| Benchmark | Kolibri | Kolibri Origin | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B | Mistral Small 4 119B-A6B |
|---|---|---|---|---|---|
| AIME 2025 | 96.9 | 81.9 | 84.6 | 91.7 | 79.8 |
| AIME 2025 (DE) | 87.5 | 73.5 | 82.9 | 85.6 | 72.3 |
| AIME 2026 | 96 | 81.5 | 91 | 90.4 | 83.1 |
| AIME 2026 (DE) | 90 | 75.2 | 84.4 | 87.5 | 78.5 |
| GPQA (diamond) | 84.3 | 68.1 | 83.4 | 78 | 74.7 |
| GPQA (diamond, DE) | 81.3 | 58.5 | 80.6 | 76.6 | 72.9 |
| AA-Omniscience Index | -32.8 | -64 | -15.3 | -36.5 | -24 |
| BrowseComp | 29.4 | 4.4 | 26.9 | 29.1 | — |
| τ³-bench banking | 38.1 | 5.7 | 10.6 | 15.5 | 5.7 |
| τ²-bench retail | 69.9 | 58.5 | 71.6 | 67.5 | 62.9 |
| τ²-bench airline | 76.7 | 58.7 | 70.7 | 72.7 | 40 |
| τ²-bench telecom | 94.7 | 67.5 | 99.1 | 68.1 | 41.5 |
| BFCL v4 overall | 61.4 | 36.4 | 67.2 | 61 | 58 |
| LiveCodeBench v6 | 85.9 | 59.2 | 82.5 | 82 | 71.2 |
| HumanEval+ | 92.7 | 76.8 | 92.8 | 94.7 | 92.8 |
| LongBench Pro | 64.5 | — | 70.8 | 62.9 | 56.4 |
| AA-LCR | 68.3 | — | 69.7 | 67 | 52.3 |
This model's scores
- AIME 202596.9%
- LiveCodeBench v685.9%
- GPQA Diamond84.3%
- GPQA Diamond (German)81.3%
- MMLU-Pro (CoT)80%
- SWE-Bench Verified66.4%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Strengths
- Open weights under Apache 2.0, built for German and English with a German-tuned tokenizer
- Only 3.46B of 78.1B parameters active per token, aimed at low serving cost
- Context validated to 1,048,576 tokens, with 262,144 native
- Selectable reasoning effort (none, low, medium, high) plus Hermes-style tool calling
- 84.3 on GPQA Diamond, 96.9 on AIME 2025 and 85.9 on LiveCodeBench v6 in Aleph Alpha's model card
- Trained to abstain when documents lack the answer: 44.0 AA-Omniscience non-hallucination rate against 14.8 for Kolibri Origin in the launch post
Best for
- Reach for it for German- and English-language assistants, document processing and drafting that must run on your own hardware.
- Reach for it for retrieval-augmented question answering over an organisation's own documents, where abstaining beats guessing.
- Reach for it for agentic tool calling and long-document work with contexts up to 262,144 tokens, or up to 1,048,576 when throughput matters less.
- Look elsewhere for image or audio input, or for languages other than German and English, which Kolibri does not target.
How to access
| Provider | Model ID |
|---|---|
| Hugging Face (weights) ↗ | Aleph-Alpha/Kolibri-1 |
| Self-hosted vLLM (aleph-alpha-inference) ↗ | Aleph-Alpha/Kolibri-1 |
FAQ
What is Kolibri?
Kolibri is an open-weight Mixture-of-Experts language model from Aleph Alpha, released on October 3, 2026. It has 78.1B total and 3.46B active parameters, focuses on German and English, supports reasoning effort levels and tool calling, and is licensed under Apache 2.0.
How long is Kolibri's context window?
Kolibri's native context length is 262,144 tokens, the length of its final long-context training phase. Aleph Alpha has validated quality and serving up to 1,048,576 tokens and recommends staying at or below 262,144 for latency-sensitive deployments and complex tasks.
What hardware does Kolibri need?
The FP8 checkpoint takes about 78 GB of memory. The model card lists a minimum of 2× A100 80 GB, 2× H100 SXM5, or one H200, B200 or B300, and recommends 2× H100 SXM5, 2× H200, or one B200 or B300.
How do I run Kolibri?
Install Aleph Alpha's aleph-alpha-inference package (a vLLM plugin, also shipped as the container image ghcr.io/aleph-alpha/aleph-alpha-inference) and run vllm serve Aleph-Alpha/Kolibri-1 with the kolibri1 reasoning and tool-call parsers. The server exposes an OpenAI-compatible API.
How does Kolibri compare with Qwen3.6-35B-A3B?
In Aleph Alpha's launch comparison, Kolibri leads Qwen3.6-35B-A3B on AIME 2025 (96.9 against 84.6), GPQA Diamond (84.3 against 83.4) and τ³-bench banking (38.1 against 10.6), while Qwen3.6-35B-A3B leads on BFCL v4 (67.2 against 61.4), τ²-bench telecom (99.1 against 94.7) and LongBench Pro (70.8 against 64.5).
What is Kolibri Origin?
Kolibri Origin is the 30.6B-total, 3.27B-active model Aleph Alpha built first to validate its training pipeline, with a 65,536-token longest trained length. It finished pre-training on June 11, 2026 and, per the launch post, had no public release.