A 78B German and English model with 3.46B active
Aleph Alpha released Kolibri on 3 October 2026. It is a mixture of experts language model built for German and English.
The 189 page tech report gives 78.1B total parameters and 3.46B active per token. The blog's opening rounds that to "3B active." Each token uses 4.4% of the parameters, the report says.
The Hugging Face card lists these serving facts:
- context of 1,048,576 tokens, with 262,144 or fewer recommended
- four reasoning effort levels, from none to high
- about 78 GB of FP8 weights
- a minimum of two A100 80 GB cards, two H100 SXM5 cards, or one H200, B200 or B300
Serving needs Aleph Alpha's vLLM plugin, shipped as a package and a container image.
Of the 50 layers, 10 use full attention. The other 40 attend to a window of 512 preceding tokens. Each token is routed to 6 of 384 experts.
A European lab bets on sovereignty
Aleph Alpha says it built the model in Germany and trained it on infrastructure in Germany and Finland. It pitches Kolibri at public administration, industrials and aerospace.
Training used 768 B200 GPUs, the blog says. Pre-training ran 20T tokens over 21 days, then 3.44T of mid-training and a long context stage. The blog puts that stage at 200B tokens and the card at 201B.
The company published a training content summary in the EU AI Act template. It says Aleph Alpha ran no web crawlers of its own and signed no commercial licensing deals for training data. The summary says synthetic data came from models including Gemma, Qwen, Mistral Nemo, GLM and Kimi K2.6.
For how the AI Act's transparency duties are landing, see EU AI Act transparency war room runs while high-risk clock moves to 2027.
Apache 2.0 for the weights only
The card says Apache 2.0 rights "only apply to the weights and configuration files" in the repository. It adds that the license "especially does not extend to underlying code, model architecture, parameter settings or any training method."
That makes the weights free to run and fine tune. Aleph Alpha keeps rights to everything around them. For an open weights release with a broader license, see Xiaomi releases MiMo-V2.6-Pro and Flash as MIT-licensed open weights with more than 7,000 RL environments.
When The Frontier checked at 14:56 SGT on 4 October, the Hugging Face repo had 278 likes. The inference plugin repo had 14 stars.
Pick it for German work on modest hardware
Aleph Alpha says it ran every benchmark on its own harnesses, at each model's highest reasoning effort. The Frontier found no independent test.
| Benchmark | Kolibri | Best rival in Aleph Alpha's table |
|---|---|---|
| Overall, English | 75.5 | Qwen3.8 27B dense, 80.2 |
| Overall, German | 70.8 | Qwen3.8 27B dense, 79.9 |
| GPQA Diamond | 84.3 | Qwen3.8 27B dense, 89.2 |
| AIME 2025, English | 96.9 | Qwen3.8 27B dense, 97.9 |
| SWE-Bench Verified | 66.4 | Qwen3.6 35B-A3B, 73.8 |
| TerminalBench 2.1 | 27.7 | Qwen3.8 27B dense, 76.8 |
Aleph Alpha's table lists Qwen3.8 27B as a dense model with 27B active parameters. The blog says Kolibri sits on the Pareto frontier of quality against serving cost, as the chart above shows. That claim measures throughput on eight B200 GPUs, per the figure caption.
Pick it when German quality, a small active size and a European supply chain matter together. The card fits it to drafting, question answering over in-house documents and research tools.
Test other models for agentic coding and terminal work. Several rivals score higher there in Aleph Alpha's own table, including Qwen3.5 35B-A3B at 39.7 on TerminalBench 2.1.
Human review, hallucination and a data mix gap
The card says Kolibri suits workflows "in which a person reviews the model's output before it is acted on." In decision support, it says the model belongs on the advisory side.
On AA-Omniscience, Kolibri abstained on 44.0% of items in place of a wrong answer, the blog says. Qwen3.6 35B-A3B's rate was 56.7%.
Aleph Alpha's own documents disagree on the German share:
- The blog says German was 21.3% of pre-training tokens, about 4.3T, against roughly 62% English and 14% code.
- The card and the training content summary give about 23.9% German, 62.5% English and 13.6% code.
The blog says the 21.3% figure counts tokens seen after upsampling a 2.4T unique German pool.
Over 21 days of pre-training, the team hit 38 unplanned interruptions. The pipeline restarted on its own each time, the blog says.
The card estimates training energy at about 950 MWh, including data centre overhead.
Aleph Alpha post, report, weights and plugin
The blog, tech report, model card, training content summary and plugin repo are listed under Sources below.
