← Back to all posts
News

Aleph Alpha Open-Sources Kolibri, a 78B German-English Model With 3.5B Active

October 3, 2026 · 22:07 UTC · News
Aleph Alpha Open-Sources Kolibri, a 78B German-English Model With 3.5B Active

TL;DR

Aleph Alpha, the Heidelberg AI company, released Kolibri on October 3 as a full open-weight model under Apache 2.0. It is a 78.1B-parameter mixture-of-experts (MoE) transformer that activates only 3.46B parameters per token, was trained on roughly 24T tokens on 768 NVIDIA B200 GPUs in Germany and Finland, and is validated out to 1,048,576 tokens of context. On Aleph Alpha's own evals it scores 75.5 overall in English and 70.8 in German, ahead of Nemotron 3 Super, Qwen3.6-35B and Mistral Small 4. The weights ship in FP8 and fit on a single H200 or B200. The catch: every number is vendor-reported, and you need Aleph Alpha's vLLM plugin to serve it.


What actually shipped

The Kolibri-1 model card on Hugging Face carries the weights and configs under Apache 2.0, with a BF16 variant and several community quantizations alongside. There is also a 189-page technical report, which is more disclosure than most closed labs give you for a press release.

The headline specs, straight from the card:

  • Size: 78,103,074,560 total parameters, 3,457,573,120 active per token (about 4.4%).
  • Shape: 50 layers, 384 routed experts plus 1 shared expert per layer, 6 experts picked per token, hidden size 2,560, 128,000-entry vocabulary.
  • Context: 262,144 tokens native, validated to 1,048,576. Aleph Alpha recommends staying at or under 262K for serving efficiency.
  • Precision: FP8 (e4m3) weights in 128x128 blocks, FP8 KV cache, with embeddings, LM head, norms and the MoE router kept in BF16.
  • Modes: reasoning effort settings of none, low, medium and high, plus Hermes-style tool calling.
  • Knowledge cutoff: June 18, 2026, for both English and German.

Kolibri is German for hummingbird, which is a fair name for something with 78B parameters that only flaps 3.5B of them at a time.

The benchmarks, with the usual asterisk

Aleph Alpha's launch post compares Kolibri against three open models and its own predecessor, Kolibri Origin (a 30.6B MoE with 3.27B active that finished pre-training on June 11 and served as the infrastructure test run). The comparison set is NVIDIA's Nemotron 3 Super (120B), Mistral Small 4 (119B) and Qwen3.6-35B.

overall score, vendor-reported (higher is better) EN DE Kolibri75.570.8 Nemotron 3 Super73.067.9 Qwen3.6-35B71.467.3 Mistral Small 463.161.4 Kolibri Origin54.146.4
With 3.46B active parameters, Kolibri edges out models with up to 120B total on Aleph Alpha's own overall scores.

A few individual numbers from the launch post:

  • AIME 2025: 96.9, versus 91.7 for Nemotron 3 Super, 84.6 for Qwen3.6-35B and 79.8 for Mistral Small 4.
  • AIME 2026: 96.0, versus 91.0, 90.4 and 83.1 for the same three.
  • GPQA Diamond: 84.3, a hair over Qwen3.6-35B at 83.4.
  • HumanEval+: 92.7, where it actually loses to Nemotron 3 Super (94.7) and roughly ties the others at 92.8.

Independent coverage from TestingCatalog also reports 85.9 on LiveCodeBench v6 and 61.4 on BFCL v4. The model card lists 63.2 for the base model on RULER at 1M tokens.

Read these the way you read any launch-day table: the vendor picked the opponents, the harness and the settings. Nothing here has been reproduced by a third party yet, and the overall English and German scores are composites Aleph Alpha defines itself. The math numbers are striking for a model this sparse; the coding numbers are good but not a sweep.

How you get 1M tokens out of 3.46B active

Two design choices carry most of the efficiency. First, the MoE routing: each token goes to 6 of 384 experts per layer, so compute per token tracks the 3.46B active count rather than the 78B total. Second, attention: according to TestingCatalog's reading of the report, only 10 of the 50 layers use full attention, while the other 40 use a 512-token sliding window.

Think of it as proofreading a novel with a ruler under the current line most of the time, then flipping back through the whole manuscript every fifth chapter to check the plot still holds. The sliding-window layers keep the KV cache small and local; the full-attention layers are what let the model retrieve something from 900K tokens ago.

The context was stretched in stages rather than all at once:

pre-train20T @ 16K mid-train3.44T @ 64K long-context201B @ 256K validated1,048,576 tok 768 B200s :: pre-training 21 days, ~392K GPU-hours, 6.4e23 FLOPs
Roughly 24T tokens total, with sequence length quadrupling at each stage before validation at 1M.

The pre-training mix on the card is 43.4% English web, 23.4% German web, 13.6% code, 11.9% English instruction and reasoning data, and 7.7% other. The launch post puts German at 21.3% of all pre-training tokens. The model card estimates training energy at about 950 MWh including data-center overhead.

The German-specific parts

The custom 128K tokenizer is built to keep German compound words intact instead of shredding Datenschutzgrundverordnung into confetti. On German web text, the launch post reports 4.90 bytes per token for Kolibri versus 4.35 for GPT-5's tokenizer and 3.59 for Qwen3-Next. More bytes per token means fewer tokens for the same document, which is cheaper inference and more of a German contract fitting in the window.

The other notable piece is grounding. Aleph Alpha trained abstention with what it calls a Merlin-Arthur procedure: the model (Arthur) gets either a helpful context provider (Merlin) or an adversary that strips the relevant evidence (Morgana), and cannot tell which one it is facing. The only stable strategy is to read the context carefully and decline when the evidence is not there.

AA-Omniscience: abstains instead of answering wrong Kolibri44% Kolibri Origin~15%
Abstention training roughly tripled how often Kolibri says "I don't know" instead of guessing.

On AA-Omniscience, Kolibri abstains instead of answering wrong on 44% of items, against about 15% for Kolibri Origin. For the regulated buyers Aleph Alpha is targeting (public administration, automotive, aerospace), a model that refuses to invent a statute is worth more than one that scores two points higher on GPQA.

Running it yourself

The weights are FP8, so the full model is roughly 78 GB. The card lists the minimum hardware as one H200, B200 or B300, or two 80 GB A100s or H100s. Serving goes through vLLM with Aleph Alpha's aleph-alpha-inference plugin, which provides the reasoning and tool-call parsers:

pip install 'aleph-alpha-inference>=1'
vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

Recommended sampling is temperature 1.0, top_p 0.97, top_k 128. With only 3.46B active parameters, decode speed should feel closer to a small dense model than a 78B one, provided you have the memory to hold all the experts. Homelab owners on 24 GB consumer cards will want to wait for the quantized variants to mature.

Why this matters

European "sovereign AI" has mostly meant procurement press releases and fine-tunes of American or Chinese base models. Kolibri is a from-scratch base model trained on European soil, released with weights, a permissive license and a long technical report. That is the version of sovereignty you can actually git clone.

It also lands in a crowded efficient-MoE tier where Qwen, NVIDIA and Mistral already compete hard. Kolibri's case rests on German quality, grounding and a clean Apache 2.0 license, not on beating everyone at everything. If your users write in German or your compliance team wants the training data to have stayed in the EU, it just became the obvious first model to evaluate.

Key Takeaways

  • Kolibri is a 78.1B-parameter MoE with 3.46B active per token, released by Aleph Alpha under Apache 2.0 on October 3, 2026.
  • It was trained on about 24T tokens using 768 B200s in Germany and Finland, with context validated to 1,048,576 tokens.
  • Vendor-reported overall scores are 75.5 English and 70.8 German, ahead of Nemotron 3 Super, Qwen3.6-35B and Mistral Small 4; no third-party reproductions yet.
  • A German-tuned tokenizer and Merlin-Arthur abstention training (44% abstain rate on AA-Omniscience) are the main differentiators.
  • FP8 weights need about 78 GB: one H200/B200 or two 80 GB cards, served via vLLM plus Aleph Alpha's plugin.

Sources: Aleph Alpha launch post, Kolibri-1 model card (Hugging Face), Kolibri technical report, aleph-alpha-inference (GitHub), TestingCatalog, Tejas Kumar, "How the sovereign German LLM works", Hacker News discussion

AIAleph AlphaKolibriOpen WeightsMixture of ExpertsSovereign AIGermanvLLM
CONSOLE
$