Aleph Alpha Open-Sources Kolibri, a 78B German-English Model With 3.5B Active
TL;DR
Aleph Alpha, the Heidelberg AI company, released Kolibri on October 3 as a full open-weight model under Apache 2.0. It is a 78.1B-parameter mixture-of-experts (MoE) transformer that activates only 3.46B parameters per token, was trained on roughly 24T tokens on 768 NVIDIA B200 GPUs in Germany and Finland, and is validated out to 1,048,576 tokens of context. On Aleph Alpha's own evals it scores 75.5 overall in English and 70.8 in German, ahead of Nemotron 3 Super, Qwen3.6-35B and Mistral Small 4. The weights ship in FP8 and fit on a single H200 or B200. The catch: every number is vendor-reported, and you need Aleph Alpha's vLLM plugin to serve it.
What actually shipped
The Kolibri-1 model card on Hugging Face carries the weights and configs under Apache 2.0, with a BF16 variant and several community quantizations alongside. There is also a 189-page technical report, which is more disclosure than most closed labs give you for a press release.
The headline specs, straight from the card:
- Size: 78,103,074,560 total parameters, 3,457,573,120 active per token (about 4.4%).
- Shape: 50 layers, 384 routed experts plus 1 shared expert per layer, 6 experts picked per token, hidden size 2,560, 128,000-entry vocabulary.
- Context: 262,144 tokens native, validated to 1,048,576. Aleph Alpha recommends staying at or under 262K for serving efficiency.
- Precision: FP8 (e4m3) weights in 128x128 blocks, FP8 KV cache, with embeddings, LM head, norms and the MoE router kept in BF16.
- Modes: reasoning effort settings of none, low, medium and high, plus Hermes-style tool calling.
- Knowledge cutoff: June 18, 2026, for both English and German.
Kolibri is German for hummingbird, which is a fair name for something with 78B parameters that only flaps 3.5B of them at a time.
The benchmarks, with the usual asterisk
Aleph Alpha's launch post compares Kolibri against three open models and its own predecessor, Kolibri Origin (a 30.6B MoE with 3.27B active that finished pre-training on June 11 and served as the infrastructure test run). The comparison set is NVIDIA's Nemotron 3 Super (120B), Mistral Small 4 (119B) and Qwen3.6-35B.
A few individual numbers from the launch post:
- AIME 2025: 96.9, versus 91.7 for Nemotron 3 Super, 84.6 for Qwen3.6-35B and 79.8 for Mistral Small 4.
- AIME 2026: 96.0, versus 91.0, 90.4 and 83.1 for the same three.
- GPQA Diamond: 84.3, a hair over Qwen3.6-35B at 83.4.
- HumanEval+: 92.7, where it actually loses to Nemotron 3 Super (94.7) and roughly ties the others at 92.8.
Independent coverage from TestingCatalog also reports 85.9 on LiveCodeBench v6 and 61.4 on BFCL v4. The model card lists 63.2 for the base model on RULER at 1M tokens.
Read these the way you read any launch-day table: the vendor picked the opponents, the harness and the settings. Nothing here has been reproduced by a third party yet, and the overall English and German scores are composites Aleph Alpha defines itself. The math numbers are striking for a model this sparse; the coding numbers are good but not a sweep.
How you get 1M tokens out of 3.46B active
Two design choices carry most of the efficiency. First, the MoE routing: each token goes to 6 of 384 experts per layer, so compute per token tracks the 3.46B active count rather than the 78B total. Second, attention: according to TestingCatalog's reading of the report, only 10 of the 50 layers use full attention, while the other 40 use a 512-token sliding window.
Think of it as proofreading a novel with a ruler under the current line most of the time, then flipping back through the whole manuscript every fifth chapter to check the plot still holds. The sliding-window layers keep the KV cache small and local; the full-attention layers are what let the model retrieve something from 900K tokens ago.
The context was stretched in stages rather than all at once:
The pre-training mix on the card is 43.4% English web, 23.4% German web, 13.6% code, 11.9% English instruction and reasoning data, and 7.7% other. The launch post puts German at 21.3% of all pre-training tokens. The model card estimates training energy at about 950 MWh including data-center overhead.
The German-specific parts
The custom 128K tokenizer is built to keep German compound words intact instead of shredding Datenschutzgrundverordnung into confetti. On German web text, the launch post reports 4.90 bytes per token for Kolibri versus 4.35 for GPT-5's tokenizer and 3.59 for Qwen3-Next. More bytes per token means fewer tokens for the same document, which is cheaper inference and more of a German contract fitting in the window.
The other notable piece is grounding. Aleph Alpha trained abstention with what it calls a Merlin-Arthur procedure: the model (Arthur) gets either a helpful context provider (Merlin) or an adversary that strips the relevant evidence (Morgana), and cannot tell which one it is facing. The only stable strategy is to read the context carefully and decline when the evidence is not there.
On AA-Omniscience, Kolibri abstains instead of answering wrong on 44% of items, against about 15% for Kolibri Origin. For the regulated buyers Aleph Alpha is targeting (public administration, automotive, aerospace), a model that refuses to invent a statute is worth more than one that scores two points higher on GPQA.
Running it yourself
The weights are FP8, so the full model is roughly 78 GB. The card lists the minimum hardware as one H200, B200 or B300, or two 80 GB A100s or H100s. Serving goes through vLLM with Aleph Alpha's aleph-alpha-inference plugin, which provides the reasoning and tool-call parsers:
pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
--reasoning-parser kolibri1 \
--tool-call-parser kolibri1 \
--enable-auto-tool-choice
Recommended sampling is temperature 1.0, top_p 0.97, top_k 128. With only 3.46B active parameters, decode speed should feel closer to a small dense model than a 78B one, provided you have the memory to hold all the experts. Homelab owners on 24 GB consumer cards will want to wait for the quantized variants to mature.
Why this matters
European "sovereign AI" has mostly meant procurement press releases and fine-tunes of American or Chinese base models. Kolibri is a from-scratch base model trained on European soil, released with weights, a permissive license and a long technical report. That is the version of sovereignty you can actually git clone.
It also lands in a crowded efficient-MoE tier where Qwen, NVIDIA and Mistral already compete hard. Kolibri's case rests on German quality, grounding and a clean Apache 2.0 license, not on beating everyone at everything. If your users write in German or your compliance team wants the training data to have stayed in the EU, it just became the obvious first model to evaluate.
Key Takeaways
- Kolibri is a 78.1B-parameter MoE with 3.46B active per token, released by Aleph Alpha under Apache 2.0 on October 3, 2026.
- It was trained on about 24T tokens using 768 B200s in Germany and Finland, with context validated to 1,048,576 tokens.
- Vendor-reported overall scores are 75.5 English and 70.8 German, ahead of Nemotron 3 Super, Qwen3.6-35B and Mistral Small 4; no third-party reproductions yet.
- A German-tuned tokenizer and Merlin-Arthur abstention training (44% abstain rate on AA-Omniscience) are the main differentiators.
- FP8 weights need about 78 GB: one H200/B200 or two 80 GB cards, served via vLLM plus Aleph Alpha's plugin.
Sources: Aleph Alpha launch post, Kolibri-1 model card (Hugging Face), Kolibri technical report, aleph-alpha-inference (GitHub), TestingCatalog, Tejas Kumar, "How the sovereign German LLM works", Hacker News discussion