Aleph Alpha’s Kolibri Is a Lightweight Sovereign LLM

Aleph Alpha has released Kolibri, an open-weight large language model for German and English. It’s a mixture-of-experts model with 78.1 billion total parameters, but it only uses about 3.46 billion of them for each token. According to Hacker News, the model came out on 3 October 2026 under the Apache 2.0 license. The weights are on Hugging Face, and Aleph Alpha trained it from scratch on infrastructure in Germany and Finland.

The name means “hummingbird” in German, which fits a model built to be light. Aleph Alpha’s own evaluation says Kolibri beats every model of its size that it was compared against, in both German and English.

🔍 What’s Inside

Here are the main specs, as detailed in Hacker News, which drew on Aleph Alpha’s 189-page technical report and the model card:

  • Parameters: 78.1B total, 3.46B active per token (4.4%)
  • Languages: German and English
  • Context window: 262,144 tokens natively, tested up to about 1 million
  • Reasoning modes: Four levels (none, low, medium, high)
  • Tool calling: Supported
  • Knowledge cutoff: 18 June 2026
  • Training: About 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 GPUs
  • Memory: About 78 GB of weights in FP8

⚙️ How It Stays Light

Kolibri has 50 layers. Each one holds 384 small sub-networks called “experts,” plus one shared expert that every token passes through. A router sends each token to just 6 of the 384. That’s why the model does the compute work of a 3.5B model while holding the knowledge of a much bigger one.

There’s a catch, though. All 384 experts have to sit in memory even when only 6 are active. The model card puts it plainly: “the full model must be held in memory even though only part of it is active at any time.” So you’ll get small-model speed, but you’ll need hardware with big-model memory.

🇩🇪 A Tokenizer Built for German

This is the part I find most interesting. German glues words together into long compounds, and tokenizers trained mostly on English chop them into pieces. The source gives a clear example: GPT-5’s tokenizer splits “Bundesverfassungsgericht” (the Federal Constitutional Court) into 6 tokens. Kolibri’s tokenizer needs only 2.

Aleph Alpha built a 128,000-token vocabulary with a new algorithm it calls UniBPE. It keeps the bottom-up merging of byte-pair encoding but scores each merge with a Unigram objective, which better respects how German builds words. The report says it uses 11.2% fewer tokens on German text than GPT-5’s tokenizer, the best result among the tokenizers it compared.

Fewer tokens matter in practice. You pay less per document, inference runs faster, and the same context window holds more real text.

🛡️ What “Sovereign” Means Here

Aleph Alpha uses the word in two ways. First, the team built and trained the model in Europe “under European and German law, with no foreign control.” Second, customers get “full freedom of deployment and intellectual-property safety.” In practice, a ministry or a car supplier can run Kolibri on its own servers, keep its data in-house, and nobody can change or switch off the model under them.

The company says it designed the model with the EU AI Act in mind “from the ground up,” and it has signed the EU’s General-Purpose AI Code of Practice.

It’s also open about the limits of that label. Some non-European models helped build the training data. Google’s Gemma 4 rephrased English web text, Mistral-NeMo handled German, and Qwen3-32B labeled data for the quality filters. Aleph Alpha says it then filtered that data for the political bias such models can carry.

⚠️ Caveats Worth Noting

  • Memory is still heavy. Expect about 78 GB for the weights alone in FP8.
  • The license only covers the weights. Apache 2.0 applies to the weights and config files. Aleph Alpha keeps the rights to its training code and methods.
  • The benchmarks are self-reported. The performance claims come from Aleph Alpha’s own evaluation, so independent tests will matter.
  • Experts aren’t topic specialists. Research on MoE models suggests experts tend to handle token patterns like punctuation or proper nouns, not subjects like “German law.”

Why It Matters

Europe has long been cast as the place that regulates while others build. Kolibri pushes back on that. It’s a competitive, permissively licensed model built around compliance, not one that had compliance bolted on later.

For German-speaking companies in regulated industries like government, automotive, finance and healthcare, the case is strong. You get efficient inference, a tokenizer that’s cheaper on German text, and full control over where the model runs.

My advice: if your workloads are mostly German or bilingual and data residency is a hard requirement, put Kolibri on your shortlist and benchmark it against your current setup. Just plan your GPU memory budget before you do. More details are available in the original report on Hacker News.

Scroll to Top