← All posts / Models

78B Parameters, 3B Active, Apache 2.0: Aleph Alpha's Kolibri Lands as Germany's Sovereign Open-Weight Bet

On the Day of German Reunification, Aleph Alpha open-sourced Kolibri: a 78B-total/3.5B-active MoE trained on 20T tokens with 21.3% organic German data, 1M-token context, and abstention training for regulated industries.

78B Parameters, 3B Active, Apache 2.0: Aleph Alpha's Kolibri Lands as Germany's Sovereign Open-Weight Bet

On October 3 — the Day of German Reunification, a date the Heidelberg company chose with intent — Aleph Alpha released Kolibri, its new sovereign language model, publishing the full weights on Hugging Face under an Apache 2.0 license. In a year when most frontier attention goes to trillion-parameter systems, Kolibri makes the opposite argument: for the public administration, industrial, and aerospace buyers it targets, the model that wins is not the biggest one, but the one a ministry can run on its own hardware, in its own language, under its own law.

The release is also the first major product to emerge from a restructured organization. Aleph Alpha is being folded into Canada’s Cohere under a roughly $20 billion combination agreement signed September 16, with $600 million backing from Schwarz Group, the Lidl-owning retail conglomerate. Founder Jonas Andrulis has moved to the board; Ilhan Scheer is now sole CEO after co-manager Reto Spörri departed in September. Kolibri is the first proof of where the new leadership is planting its flag — not in the scale race against OpenAI or Anthropic, but in the compliance-heavy, infrastructure-sensitive corner of the market that Schwarz Group’s enterprise relationships already touch.

What Kolibri actually is

Kolibri is an English-German Mixture-of-Experts Transformer with 78.1 billion total parameters and roughly 3.5 billion active per token. It supports context lengths of up to 1 million tokens and was trained on approximately 20 trillion tokens of pre-training data. Two variants are downloadable — Kolibri-1 and a BF16 edition — and both ship with vLLM support out of the box.

The architecture is where the engineering discipline shows. Kolibri uses 384 small experts rather than fewer wide ones (Aleph Alpha’s tests found the finer-grained split performs better), with 6 experts active per token and one shared expert across all 50 layers. Only 10 of those layers run full attention; the other 40 use a tight 512-token sliding window, which keeps decode compute and memory bounded regardless of context length. For routing between experts during training, the team introduced exact quantile balancing — an exact computation of what Kimi K3 previously estimated from histograms — and reports that the exactness improves both load balance and final model quality.

The whole run was trained on 768 B200 GPUs in three stages: 20T tokens at 16k sequence length over 21 days, then 3.44T tokens of mid-training at 64k, then 200B tokens of long-context adaptation at 256k — nearly 24T tokens in total. Over the 21-day pre-training stretch the job hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, all handled automatically by the training pipeline without a person stepping in; the cluster restarted on different nodes and resumed from a checkpoint at most 250 steps back.

A model that is German by design, not by translation

The most distinctive part of the project is the data story. Aleph Alpha needed roughly 4T German tokens to hit its ~20% German target, and after deduplication the open German web offered just 390B. Rather than lean on machine translation — which, as the team argued in its earlier “Sauerkraut, Not Burgers” work, produces “translationese” and imports the cultural fingerprints of the English web — they closed the gap three ways.

First, a German-specialized Common Crawl pipeline: standard English filters quietly delete German administrative prose, because German words routinely exceed English mean-length bounds, so the team retuned the filtering parameters and recovered 1.3T unique tokens of organic German web text. Second, rephrasing: an LLM rewrote existing German documents as encyclopedia entries, dialogues, or passages — same facts, new surface forms, still culturally German (“chancellor, not president,” as the blog puts it), yielding about 1T tokens. Translation itself contributed only 6% overall, and none in the final Kolibri run. The result: German entered training as a 2.4T-token unique pool, 80% of it curated or generated in-house, with the model ultimately seeing German at 21.3% of pre-training tokens.

The team also built a bilingual English-German tokenizer using a new method they call UniBPE, which combines BPE’s bottom-up merging with the Unigram training objective. On German web text, Kolibri’s 128k-vocabulary tokenizer compresses at 4.90 bytes per token — better than GPT-5 (4.35), DeepSeek V4 (3.72), Kimi K3 (3.28), and GLM 5.3 (3.93) in their comparison — and it splits German compounds along morpheme boundaries that competitors cut straight through: Bundessozialgericht becomes Bundes+sozial+gericht+es rather than a mess of sub-lexical fragments. Better compression is not cosmetic; fewer tokens per task means cheaper, faster inference.

Benchmarks: small active parameters, frontier-adjacent results

Against open-weight peers — Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B, Mistral Small 4 119B-A6B — Kolibri posts strong numbers: 96.9 on AIME 2025, 96.0 on AIME 2026 (and 90.0 on the German edition), 84.3 on GPQA Diamond (81.3 in German), 85.9 on LiveCodeBench v6, and 92.7 on HumanEval+. Aleph Alpha claims Kolibri matches models with up to four times its active parameter count and sits on the Pareto frontier for quality versus serving cost in both languages. The economics explain the sizing: a 123B variant they tested could juggle only 3 concurrent 256k-token queries on two H100s, while the 78B Kolibri handles 18 and decodes 28% faster.

For hallucination-sensitive deployments, Kolibri was trained with abstention data and Aleph Alpha’s Merlin-Arthur protocol — a three-player game in which “Arthur” (the shipped model) must answer from real contexts supplied by “Merlin” but abstain on redacted contexts crafted by “Morgana,” making any guess, even a lucky one, incorrect. Kolibri abstains instead of answering wrong on 44% of AA-Omniscience items (its predecessor: 14.8%), and posts 0.23 on Aleph Alpha’s own M/A grounding score where most compared models score 0. The model also reasons at four selectable effort levels — none, low, medium, high — letting customers trade compute against answer quality per request.

Why sovereignty is the product

Every design choice loops back to the same pitch: control. Kolibri was built in Germany, trained on infrastructure in Germany and Finland, under European and German law, with what the company describes as no foreign control over the pipeline from data curation through post-training. It was engineered with the EU AI Act, the General-Purpose AI Code of Practice, and GDPR in mind. Customers get full deployment freedom — on-premise, air-gapped, wherever — and because the model is small-active-parameter, it runs efficiently without shipping internal documents to a third-party inference cloud.

The strategic read: this is Mistral’s French playbook, narrowed to the DACH slice that must answer to German data-protection law and cannot route workloads through a hyperscaler in Virginia. Whether German agencies actually standardize on Kolibri — and whether it holds its own against the open-weight Chinese models German enterprises quietly test — is now a question of customers, not press releases. But as a statement of what “sovereign AI” means in practice, Kolibri is unusually concrete: open weights, a documented supply chain, a bilingual-by-design tokenizer, and a model trained to say “I don’t know.”