Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,678 words · 5 segments analyzed
Kolibri is an open-weight large language model (LLM) from Aleph Alpha for German and English: a mixture of experts with 78 billion parameters that only uses about 3.5 billion of them for each token it reads or writes. It came out on 3 October 2026 under the Apache 2.0 license, the weights are on Hugging Face, and it was trained from scratch on infrastructure in Germany and Finland. (Kolibri is German for hummingbird, which is cute for a model whose whole trick is being light.) I live in Germany, and at SmashingConf New York in 2024 I told the room what I’d heard in the US when I said where I’m based: “you regulate, you don’t innovate.” It hurt to hear, and what I wished for on that stage was the middle, “the right balance between innovation and regulation around data privacy, data stewardship, environmental constraints and energy requirements.” Kolibri is a pretty direct answer to that: a German team built it with the EU AI Act in mind “from the ground up”, and in Aleph Alpha’s own evaluation it scores above every compared model of its size in both languages.
Huge congrats to everyone at Aleph Alpha who built it, my good friend Michael Hofmann among them! This post is about how Kolibri works, where it’s strong, where it isn’t, how to run it, and when it’s the right pick. Everything here comes from Aleph Alpha’s 189 page technical report, the model card and their launch post, plus one experiment I ran on its tokenizer. What is Kolibri? Kolibri 1 Parameters 78.1 billion in total, 3.46 billion per token (4.4%) Languages German and English Context 262,144 tokens natively, tested up to 1,048,576 License Apache 2.0 for the weights and configuration files (Aleph Alpha keeps the rights to its training code and methods) Memory about 78 GB of weights in 8-bit floating point (FP8) Reasoning 4 levels: none, low, medium and high Tool calling Yes Knowledge cutoff 18 June 2026 Training about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 graphics processing units (GPUs) Aleph Alpha calls Kolibri sovereign, and in their launch post that means 2 things. The first is how it was built: “teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control.” The second is what customers get: “full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property.” In plain words, a ministry or a car supplier can run it on its own servers, with its data never leaving the building, and nobody can change or switch off the model under them. Aleph Alpha has also signed the European Union’s General-Purpose AI (GPAI) Code of Practice. Sovereign doesn’t mean that nothing from outside Europe went in, and the model card says so itself: English web text was rephrased with Google’s Gemma 4, German with Mistral-NeMo, and Qwen3-32B labeled data for the quality filters. They then filtered the training data for the political bias such models can have, which they’ve measured in Chinese open models themselves. How Kolibri works Kolibri is 6 ideas stacked on top of each other, and each one is there to make German cheaper, longer or more honest. 1. 384 specialists, and each token sees 6 In a normal (dense) model, every token goes through every parameter. In a mixture of experts (MoE), each layer has a crowd of small sub-networks called experts and a router that picks a few of them for each token. Kolibri has 50 layers, each with 384 experts plus 1 shared expert that every token goes through, and its router sends each token to 6 of the 384. That’s how 78.1 billion parameters turn into 3.46 billion of actual work per token.
In 2024 I gave a talk called Why Small Language Models are the future, and I argued for “smaller language models with fewer parameters and fewer places things can go wrong that require lesser compute.” My analogy was a doctor who has read every medical book in the world against a specialist in hematology: go to the first one with a blood condition and “they may not get it right cuz they know too much.” A mixture of experts puts a hospital full of specialists inside one model, and the router is the receptionist who sends each token to the right 6. The analogy breaks in 2 places though.
The experts aren’t neat topics like “German law”: when researchers look inside MoE models they mostly find experts for patterns of tokens, like punctuation or proper nouns, not subjects a person would pick. And the hospital has to keep all 384 specialists on staff even if you only see 6, so Kolibri computes like a 3.5 billion parameter model but needs the memory of a 78 billion parameter one. The model card says it plainly: “the full model must be held in memory even though only part of it is active at any time.” 2. A tokenizer that reads long German words A model doesn’t read letters or words, it reads tokens: chunks of text from a fixed vocabulary, picked when the tokenizer is trained. German glues words together into long compound words, and a tokenizer that learned mostly from English chops them into pieces.
Here’s the German name of the Federal Constitutional Court, split by the tokenizer GPT-4o and GPT-5 use (o200k_base, through OpenAI’s tiktoken), and by Kolibri’s: o200k_base (GPT-5): Bund | es | ver | fass | ungs | gericht 6 tokens Kolibri: Bundes | verfassungsgericht 2 tokens Kolibri’s tokenizer has 128,000 tokens, trained with a new algorithm Aleph Alpha calls UniBPE: it keeps the bottom-up merging of byte-pair encoding (BPE) and picks each merge with a different scoring rule (the Unigram objective), which respects how German builds words. The report says it needs 11.2% fewer tokens for German text than GPT-5’s tokenizer, the best of the 9 others they measured. I wanted to see that for myself, so I ran 6 tokenizers over all of the Basic Law for the Federal Republic of Germany, the German constitution (185 KB of very German legal text), and over its official English translation: Tokenizer German tokens More than Kolibri English tokens More than Kolibri Kolibri 1 35,190 39,875 o200k_base (GPT-4o, GPT-5, GPT-OSS) 41,482 17.9% 39,737 -0.3% Qwen3.5 35B-A3B 42,907 21.9% 41,650 4.5% Mistral Small 4 43,478 23.6% 41,301 3.6% Gemma 4 43,850 24.6% 41,564 4.2% On legal German, Kolibri needed 15% fewer tokens than GPT-5’s tokenizer, even more than Aleph Alpha’s own 11.2%, and in English it tied with it. Wild. Fewer tokens means fewer steps to read or write the same German text, and more German fits in the same context window. (I counted each one with its own tokenizer.json through Hugging Face’s tokenizers library, except o200k_base, which I counted with tiktoken.) 3. Most layers only look nearby 40 of Kolibri’s 50 layers use sliding-window attention: each token only looks at the 512 tokens before it. Every 5th layer looks at everything before it. It’s like reading a long contract while mostly paying attention to the sentence you’re on, and every few pages stopping to think about all of it, and it’s what keeps a 1 million token context affordable. There’s a clever detail in there too. Only the sliding-window layers know where a token sits (through rotary position embeddings), and the full-attention layers don’t, so the context stretches past the 262,144 tokens it was trained on without any extra position tricks. Aleph Alpha validated it up to 1,048,576. In my talk Unlocking Value with AI Today I called finite context one of “the big three” problems of generative AI, next to hallucination and the knowledge cutoff, and Kolibri goes after all 3. At 1 million tokens, on the RULER long-context test, Kolibri’s base model scores 63.2, against 57.5 for Qwen3.5 35B-A3B’s base model. 4. It thinks in German Reasoning models think before they answer, and even on German prompts they mostly think in English. Aleph Alpha posted about this on 24 September and wrote it up as Through the Valley of Tears: they generated about 800,000 German reasoning examples, and found that a little German reasoning data is worse than none. Their model’s German math score dropped from 70.2 to 48.3, because its German thoughts kept going around in circles and never finished, and it only climbed back (to 67.3) with a lot more German data. Kolibri got the lot more. It reasons in German on German prompts, and its German math scores are the best of the models with about 3 billion active parameters: 87.5 on the American Invitational Mathematics Examination (AIME) 2025 in German, against 84.4 for the next best, NVIDIA’s Nemotron 3 Nano. 5. It’s trained to say “I don’t know” Hallucination was number 1 of my big three, and the fix in that talk was retrieval augmented generation (RAG): you look up good, authoritative information and put it in the prompt. RAG only works if the model admits when the documents don’t have the answer, though, and that’s what Aleph Alpha trained for, with their own method, the Merlin-Arthur protocol. It works like a game with 3 players. Arthur is the model, and he gets a question with parts of the supporting document hidden. Merlin hides parts so that the correct answer gets easier to find, and Arthur is trained to answer those. Morgana hides the evidence the answer depends on, and Arthur is trained to say he doesn’t know. Arthur never knows which of the 2 he’s facing, so the only way to win is to actually check whether the evidence in front of him supports an answer. It shows. On Artificial Analysis’s Omniscience test, when Kolibri didn’t know an answer, it said so (or gave a partial answer) 44% of the time instead of making one up. Qwen3.5 35B-A3B did that 11.1% of the time and GPT-OSS 120B 23.7%. Of all the mixture-of-experts models Aleph Alpha compared, only Qwen3.6 35B-A3B did better, at 56.7%. 6.