Aleph Alpha's sovereign, open-weight German-and-English language model — a mixture-of-experts model you self-host, with a long context window and selectable reasoning levels.
Last verified · 2026-10-03 · by Moe Ameen
Kolibri is an open-weight language model from Aleph Alpha, the Heidelberg-based lab, released on October 3, 2026 under the permissive Apache 2.0 license. The name is German for "hummingbird," a nod to its efficient design. The weights live on Hugging Face as Aleph-Alpha/Kolibri-1, so you can download, run, and fine-tune it yourself rather than renting it through an API.
Technically it is a mixture-of-experts (MoE) model: 78 billion total parameters, but only about 3.46 billion active per token, routed across 384 experts in each of its 50 layers — the design that keeps inference cheap for the model's size. It handles a native context of 262,144 tokens, extendable toward roughly one million, and exposes four reasoning effort levels (none, low, medium, and high) so you can trade speed for depth per request. Aleph Alpha says it trained on about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 GPUs in Germany and Finland, with a knowledge cutoff of June 18, 2026.
The positioning is "sovereign AI": built in Europe under European law, aligned to the EU AI Act, the General-Purpose AI Code of Practice, and GDPR, and designed for on-premise deployment so data stays in-house. It is built for German-first, regulated, mission-critical work and reasons in German directly rather than translating through English.
Be clear about the scope and the floor. Kolibri is text-only — no images, audio, or video — and while it leads on raw code-generation benchmarks like LiveCodeBench, it trails newer open-weight competitors on agentic coding tasks (SWE-bench, Terminal-Bench) and multi-turn tool calling. Running it is not light either: Aleph Alpha lists a minimum of roughly 78 GB of GPU memory (two 80 GB A100s or H100s, or a single H200, B200, or B300), and at launch no third party hosted it, so self-hosting was the only way in.
Kolibri's standout trait for a content workflow is its appetite for context: you can feed it a genuinely large pile of source material — a back catalogue of transcripts, a research report, a set of policy documents, a quarter of meeting notes — inside its 262,144-token window and have it draft grounded German or English copy from the whole thing at once. That is the hard, unglamorous part of content: reading the raw material and turning it into a coherent draft. What Kolibri does not do is anything after the draft — it outputs text and stops, makes no media, and publishes nothing.
That is the exact handoff [Kompozy](/) is built for. Take the drafts Kolibri produces from your long-context source and bring them in as source material, then let Kompozy spin one week's worth of raw text into a full content batch: [persona and avatar shorts](/glossary/persona-shorts) and HeyGen video from a script, [Clipped Shorts](/glossary/clipped-short) if you also have long-form footage, brand-exact carousels and quote cards for the strongest lines, and a [blog article plus an email newsletter](/glossary/output-buckets) from the long-form draft. A [Persona Brief](/glossary/persona-brief) holds one voice and your banned words across all of it, and [autopilot](/glossary/autopilot) schedules and publishes the batch across the eight social platforms plus blog and email through a per-post review step. Kolibri reads the pile and writes; Kompozy turns the writing into finished, scheduled content. Kompozy runs Claude and OpenAI for its own generation and does not host Kolibri — the two meet at the draft.
Kolibri is an open-weight, text-only language model from the Heidelberg lab Aleph Alpha, released October 3, 2026 under Apache 2.0. It is a mixture-of-experts model with 78 billion total parameters (about 3.46 billion active per token), built for German and English, with a context window up to roughly one million tokens and four selectable reasoning levels. It is positioned as sovereign AI — built in Europe, aligned to the EU AI Act and GDPR, and designed to run on your own infrastructure.
The weights are free to download under the permissive Apache 2.0 license, so there is no per-token API fee. Running it is not light, though: Aleph Alpha lists a minimum of roughly 78 GB of GPU memory — two 80 GB A100s or H100s, or a single H200, B200, or B300 — and at launch no hosted provider served it, so self-hosting was the only route.
Yes — that is its clearest strength. It is trained to reason in German directly rather than translating through English, more than a fifth of its training data was German, and it posts notably strong German reasoning and math results. For teams producing German copy, scripts, or long-form text, that native-German quality is the main reason to pick it over a model that treats German as a secondary language.
No. Kolibri is text-only — it drafts and reasons over text and generates no images, video, captions, or finished posts, and it does not publish anywhere. To turn its German or English drafts into finished, multi-platform content, pair it with Kompozy, which generates persona and avatar video, clips, carousels, quote cards, blogs, and newsletters and publishes them across the eight social platforms plus blog and email.