The Heidelberg lab shipped Kolibri on the Day of German Reunification — a 78B-parameter mixture-of-experts model, open-weighted under Apache 2.0, that reasons in German natively and is built to run on your own infrastructure.
2026-10-03 · by Moe Ameen
On October 3, 2026 — the Day of German Reunification — the Heidelberg-based lab Aleph Alpha released Kolibri, an open-weight language model built for German and English. The name is German for "hummingbird," a nod to its efficient design. The weights are published on Hugging Face as Aleph-Alpha/Kolibri-1 under the permissive Apache 2.0 license, so anyone can download, run, and fine-tune it.
Kolibri is a mixture-of-experts (MoE) model: 78 billion total parameters, but only about 3.46 billion active per token, routed across 384 experts in each of its 50 layers. It supports a native context of 262,144 tokens, extendable toward roughly one million, and offers four reasoning effort levels — none, low, medium, and high — so you can trade speed for depth per request. Aleph Alpha says it was trained on about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 GPUs in Germany and Finland, with a knowledge cutoff of June 18, 2026.
The pitch is "sovereign AI." Aleph Alpha built the model in Europe under European law, with the EU AI Act, the General-Purpose AI Code of Practice, and GDPR in mind, and designed it to be deployed on-premise so an organization's data never leaves its own infrastructure. It is aimed squarely at regulated, mission-critical work — public administration, industrials, aerospace — and it is trained to reason in German directly rather than translating through English the way many large models do.
Kolibri is text-only, not multimodal, and has real trade-offs: it leads on raw code-generation benchmarks like LiveCodeBench but trails newer open-weight competitors on agentic coding tasks and multi-turn tool calling, and it leans on retrieval rather than recalling facts from memory — it is trained to abstain when it does not know rather than fabricate. Running it is not trivial either: it needs roughly 78 GB of GPU memory (two 80 GB A100s or H100s, or a single H200, B200, or B300), and at launch no hosted provider served it, so self-hosting was the only way to try it.
Kolibri is a drafting engine, not a content engine, and that distinction is the whole opportunity. If you serve a German-speaking audience, the honest workflow is to draft your source text where the language quality and data control live, then produce and publish the finished content somewhere built for it. Run Kolibri on your own hardware to write a German-language script, an article outline, or a set of campaign notes — on infrastructure you control, which is the entire point of a sovereign model — and treat that text as raw material. [Kompozy](/) is where it becomes a campaign: bring the draft in as a source and it fans the idea into formats a language model cannot make — [persona and avatar shorts](/glossary/persona-shorts), brand-exact carousels, quote cards, [blog articles and email newsletters](/glossary/output-buckets), and platform-native text — each held to one voice by a [Persona Brief](/glossary/persona-brief) and your banned-word list.
Kompozy does not run Kolibri, and it should not pretend to — its own generation uses Claude and OpenAI. What it adds is the half a raw model leaves undone: turning text into finished video, image, blog, and newsletter assets and shipping them on a schedule across the eight social platforms plus blog and email through a per-post review pipeline with [autopilot](/glossary/autopilot). Draft with a model you govern; publish everywhere from one queue.
Kolibri is an open-weight, text-only AI language model from the Heidelberg lab Aleph Alpha, released October 3, 2026 under the Apache 2.0 license. It is a mixture-of-experts model with 78 billion total parameters (about 3.46 billion active per token), built for German and English and positioned as "sovereign AI" — built in Europe, aligned to the EU AI Act and GDPR, and designed to run on your own infrastructure.
Aleph Alpha released Kolibri on October 3, 2026 — the Day of German Reunification — with open weights on Hugging Face as Aleph-Alpha/Kolibri-1 under the permissive Apache 2.0 license, meaning it can be downloaded, run, and fine-tuned freely.
The weights are free to download under Apache 2.0, so there is no per-token API fee. But it is not light to run: Aleph Alpha lists a minimum of roughly 78 GB of GPU memory — two 80 GB A100s or H100s, or a single H200, B200, or B300 — and at launch no hosted provider served it, so self-hosting was the only path.
No. Kolibri is text-only — it drafts and reasons over text and does not generate images, video, captions, or finished posts, nor does it publish anywhere. To turn its German or English drafts into finished, multi-platform content you pair it with a content engine like Kompozy, which generates video, images, carousels, blogs, and newsletters and publishes them across platforms.