Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace.
Is it deployable? Yes. The FP8 checkpoint is about 78GB and runs on a single B200, B300 or H200, or on 2 H100 SXM5 GPUs, served through vLLM with dedicated Kolibri reasoning and tool-call parsers.
What is Kolibri?
Kolibri (Kolibri-1) is a bilingual English-German MoE transformer developed end to end by teams in Germany. According to the technical research report, Aleph Alpha team controlled the full pipeline: data, architecture, training infrastructure, post-training and evaluation. Training ran on infrastructure in Germany and Finland. The design targets the EU General-Purpose AI Code of Practice, the EU AI Act and GDPR. Aleph Alpha is a signatory of that Code, and its data pipeline redacts personal data before training.
Architecture: Sparse Experts and Hybrid Attention
Kolibri stacks 50 transformer blocks with a model width of 2,560. Every MoE layer scores all 384 routed experts with a sigmoid router, sends each token to the top 6, and always runs 1 shared expert. Expert load is balanced with Exact Quantile Balancing and Load-Error Injection.
Attention uses grouped-query attention with 48 query heads and 4 KV heads. Every fifth block uses full attention without positional encoding. The other 40 blocks use sliding-window attention over the 512 preceding tokens, with RoPE. Sliding-window layers hold a fixed-size KV cache, so only 10 layers grow with context length. At matched compute, Aleph Alpha team reports the hybrid supports sequences 4 times longer than a full-attention model.
A Tokenizer Built for German
The 128,000-token vocabulary is trained with UniBPE, which builds merges like BPE but scores each merge by Unigram loss. On German text it reaches 4.90 bytes per token, versus 4.35 for the GPT-5 tokenizer. That means 11.2% fewer tokens on German web text. In English, Kolibri reaches 4.58 bytes per token against 4.67 for GPT-5.
Training: 24T Tokens, Then SFT and RL
Pre-training covered 20T tokens on 768 NVIDIA B200 GPUs, followed by 3.44T mid-training tokens at 65,536 sequence length. A 201B-token long-context stage then trained on 262,144-token sequences. Aleph Alpha added more than 2T German tokens it curated from the web or generated synthetically. Post-training combined supervised fine-tuning, mixed with MergeMix, with reinforcement learning on more than 1.2M internal tasks. The Merlin-Arthur protocol trains the model to abstain when retrieved context does not support an answer.
Interactive Explainer: How Kolibri Works
Benchmarks
Aleph Alpha team evaluated every model with the same eval-framework setup. In English, Kolibri leads GPQA Diamond (84.3), AIME 2025 (96.9) and AIME 2026 (96.0). It ties Qwen3.5 35B-A3B on the English agentic average at 63.4. It trails on BFCL v4, scoring 61.4 against 70.5 for Qwen3.5. The dense Qwen3.8 27B scores higher overall (80.2 EN, 79.9 DE), but activates about 8 times more parameters per token. Against its internal predecessor Kolibri Origin, Kolibri decodes about 2.7 times more text per GPU while scoring 21.4 points higher in English.
Kolibri vs Closest Competitors
| Feature | Kolibri-1 | Qwen3.6 35B-A3B | Nemotron 3 Super | Mistral Small 4 |
|---|---|---|---|---|
| Developer | Aleph Alpha (Germany) | Alibaba Qwen | NVIDIA | Mistral AI (France) |
| Total / active params | 78.1B / 3.46B | ~35B / 3B | 120B / 12B | 119B / 6.5B |
| Architecture | MoE, sliding-window + full attention | MoE, Gated DeltaNet hybrid | Mamba-2 + MoE + attention | MoE |
| Max context | 1,048,576 tokens | 262,144 native, ~1M with YaRN | 1M tokens | 256k tokens |
| Reasoning control | none / low / medium / high | Thinking on / off | On / off, low-effort mode | none / high |
| License | Apache 2.0 | Apache 2.0 | NVIDIA Nemotron Open Model License | Apache 2.0 |
| Overall score (EN / DE) | 75.5 / 70.8 | 71.4 / 67.3 | 73.0 / 67.9 | 63.1 / 61.4 |
| Agentic avg (EN) | 63.4 | 62.1 | 54.9 | 40.7 |
| Industry RAG avg (DE) | 67.5 | 65.8 | 57.6 | 53.4 |
| Code avg (EN) | 89.3 | 87.7 | 88.3 | 82.0 |
Specs: model cards for Kolibri-1, Qwen3.6-35B-A3B, Nemotron 3 Super and Mistral Small 4.
How to Deploy Kolibri
Install the aleph-alpha-inference package, then serve with vLLM:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 --tool-call-parser kolibri1 \
--enable-auto-tool-choiceThe default context is 262,144 tokens. Pass --max-model-len 1048576 with a max_position_embeddings override for the full 1M window. Set reasoning_effort through chat_template_kwargs. Aleph Alpha recommends temperature 1.0, top-p 0.97 and top-k 128.
Key Takeaways
- 78.1B total, 3.46B active: each token uses 6 of 384 routed experts plus 1 shared expert.
- 40 sliding-window and 10 full-attention layers keep a 1M-token context affordable.
- Trained on 24T tokens, with German above 20% of the mix.
- Top Overall score in English (75.5) and German (70.8) among the 12 MoE models Aleph Alpha compared.
- Apache 2.0 FP8 weights, serveable on a single B200 or H200 via vLLM.
Check out the Model Weights and Technical Report. All credit goes to the researcher of this project. Also,ย feel free to follow us onย Twitterย and donโt forget to join ourย 150k+ML SubRedditย and Subscribe toย our Newsletter. Wait! are you on telegram?ย now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.







