// AI TOOLS · KOLIBRI

Kolibri

Aleph Alpha's sovereign, open-weight German-and-English language model — a mixture-of-experts model you self-host, with a long context window and selectable reasoning levels.

Last verified · 2026-10-03 · by Moe Ameen

What Kolibri is

Kolibri is an open-weight language model from Aleph Alpha, the Heidelberg-based lab, released on October 3, 2026 under the permissive Apache 2.0 license. The name is German for "hummingbird," a nod to its efficient design. The weights live on Hugging Face as Aleph-Alpha/Kolibri-1, so you can download, run, and fine-tune it yourself rather than renting it through an API.

Technically it is a mixture-of-experts (MoE) model: 78 billion total parameters, but only about 3.46 billion active per token, routed across 384 experts in each of its 50 layers — the design that keeps inference cheap for the model's size. It handles a native context of 262,144 tokens, extendable toward roughly one million, and exposes four reasoning effort levels (none, low, medium, and high) so you can trade speed for depth per request. Aleph Alpha says it trained on about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 GPUs in Germany and Finland, with a knowledge cutoff of June 18, 2026.

The positioning is "sovereign AI": built in Europe under European law, aligned to the EU AI Act, the General-Purpose AI Code of Practice, and GDPR, and designed for on-premise deployment so data stays in-house. It is built for German-first, regulated, mission-critical work and reasons in German directly rather than translating through English.

Be clear about the scope and the floor. Kolibri is text-only — no images, audio, or video — and while it leads on raw code-generation benchmarks like LiveCodeBench, it trails newer open-weight competitors on agentic coding tasks (SWE-bench, Terminal-Bench) and multi-turn tool calling. Running it is not light either: Aleph Alpha lists a minimum of roughly 78 GB of GPU memory (two 80 GB A100s or H100s, or a single H200, B200, or B300), and at launch no third party hosted it, so self-hosting was the only way in.

What you can make with it

  • German-language drafts — scripts, articles, newsletters, and posts — written by a model trained to reason in German rather than translate it
  • English and bilingual copy from the same model, with selectable reasoning effort for quick drafts or careful long-form
  • Summaries, outlines, and analysis over very long source material (up to 262,144 tokens natively, extendable toward ~1M)
  • Grounded answers that abstain when the model is unsure rather than fabricating, useful for higher-stakes drafting
  • Structured extraction and rewriting of documents entirely on your own infrastructure, with no data sent to an external API
  • Fine-tuned, self-hosted variants adapted to your domain or house style, since the weights are open under Apache 2.0

How Kompozy turns Kolibri output into content

Kolibri's standout trait for a content workflow is its appetite for context: you can feed it a genuinely large pile of source material — a back catalogue of transcripts, a research report, a set of policy documents, a quarter of meeting notes — inside its 262,144-token window and have it draft grounded German or English copy from the whole thing at once. That is the hard, unglamorous part of content: reading the raw material and turning it into a coherent draft. What Kolibri does not do is anything after the draft — it outputs text and stops, makes no media, and publishes nothing.

That is the exact handoff [Kompozy](/) is built for. Take the drafts Kolibri produces from your long-context source and bring them in as source material, then let Kompozy spin one week's worth of raw text into a full content batch: [persona and avatar shorts](/glossary/persona-shorts) and HeyGen video from a script, [Clipped Shorts](/glossary/clipped-short) if you also have long-form footage, brand-exact carousels and quote cards for the strongest lines, and a [blog article plus an email newsletter](/glossary/output-buckets) from the long-form draft. A [Persona Brief](/glossary/persona-brief) holds one voice and your banned words across all of it, and [autopilot](/glossary/autopilot) schedules and publishes the batch across the eight social platforms plus blog and email through a per-post review step. Kolibri reads the pile and writes; Kompozy turns the writing into finished, scheduled content. Kompozy runs Claude and OpenAI for its own generation and does not host Kolibri — the two meet at the draft.

  1. Self-host Kolibri (or run it however you access the weights) and feed it your long-context source material — transcripts, reports, notes — to draft a grounded German or English script, article, and newsletter in one pass.
  2. Bring those drafts into Kompozy as a source, and set a Persona Brief so every output holds one voice and your banned words.
  3. Generate short-form from the script: Persona Shorts or a HeyGen avatar video, auto-captioned and reframed to 9:16, 1:1, and 16:9, plus Clipped Shorts if you have long-form footage.
  4. Turn the strongest lines into a brand-exact Carousel and Quote Graphics, and the long-form draft into a Blog Article and an Email Newsletter.
  5. Schedule and publish the whole batch across the eight social platforms plus blog and email from one queue with autopilot and a per-post review pass.

Frequently asked questions

What is Kolibri?

Kolibri is an open-weight, text-only language model from the Heidelberg lab Aleph Alpha, released October 3, 2026 under Apache 2.0. It is a mixture-of-experts model with 78 billion total parameters (about 3.46 billion active per token), built for German and English, with a context window up to roughly one million tokens and four selectable reasoning levels. It is positioned as sovereign AI — built in Europe, aligned to the EU AI Act and GDPR, and designed to run on your own infrastructure.

Is Kolibri free, and what does it take to run?

The weights are free to download under the permissive Apache 2.0 license, so there is no per-token API fee. Running it is not light, though: Aleph Alpha lists a minimum of roughly 78 GB of GPU memory — two 80 GB A100s or H100s, or a single H200, B200, or B300 — and at launch no hosted provider served it, so self-hosting was the only route.

Is Kolibri good for German-language content?

Yes — that is its clearest strength. It is trained to reason in German directly rather than translating through English, more than a fifth of its training data was German, and it posts notably strong German reasoning and math results. For teams producing German copy, scripts, or long-form text, that native-German quality is the main reason to pick it over a model that treats German as a secondary language.

Can Kolibri create images, video, or social posts?

No. Kolibri is text-only — it drafts and reasons over text and generates no images, video, captions, or finished posts, and it does not publish anywhere. To turn its German or English drafts into finished, multi-platform content, pair it with Kompozy, which generates persona and avatar video, clips, carousels, quote cards, blogs, and newsletters and publishes them across the eight social platforms plus blog and email.

Related tools

  • DeepSeek V4.1 Flash — DeepSeek's re-architected, natively multimodal Flash model — it reads images alongside text at V4-Flash pricing, and DeepSeek says it surpasses the larger V4-Pro on performance, cost, and speed.
  • Writer Palmyra X6 — Writer's enterprise flagship agentic model, launched August 13, 2026 as a post-trained variation of the open-source GLM-5.2 — built to run governed, multi-step business tasks at a lower token cost alongside a rebuilt Agent harness.
  • Claude Opus 5.5 — Anthropic's September 2026 frontier Claude model — about 20% cheaper per token than Opus 5, more than 30% faster, with a #1 debut on the Artificial Analysis Intelligence Index and a new class of safeguards. A text-output model, not an image, audio, or video generator.
  • K2 Horizon — MBZUAI's Institute of Foundation Models: a family of six fully open text models from 0.9B to 375B, released with weights, code, and training data under Apache 2.0.
  • GPT-6 Sol and Luna — OpenAI's cost-optimized GPT-6 models — Sol for complex coding and knowledge work, Luna for cheap high-volume tasks — launched September 22, 2026 with fewer mistakes at roughly half the price of the GPT-5.6 tier.

← All AI tools · Get started →