Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha
Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,358 words · 1 segments analyzed
Research Aleph Alpha 03/10/2026 On the Day of German Reunification, we are releasing our new model: Kolibri. Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active. It supports context lengths of up to 1M tokens. The model can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms. Kolibri is the result of continuous iteration of our model training effort. We first built a model training pipeline and validated it by building Kolibri Origin, a 30B total, 3B active model with a much shorter 65k token context window. Kolibri ran through the same pipeline: from data ingestion and curation, through ablations, pre-training, and post-training, to the final evals. It enabled running hundreds of ablation experiments and stable pre-training that ran without a person having to step in when hardware failed or a data connection dropped. We continuously monitored training metrics and standardized monitoring for custom benchmarks. The time we put into building and iterating on this pipeline was a valuable investment. We see it in how much better Kolibri is than Kolibri Origin, and in how little time separates their releases. Kolibri is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace. We specialized Kolibri for German, reasoning, math, agentic behavior, and further capabilities our customers need in production. The aim of this specialization was to optimize performance in our customers' specific use cases. Through specialization, customers achieve contextualized performance in their AI operations and they can monitor its economic impact, so that ROI stays measurable and grows over time. Specialization alone is not enough. Sovereignty is just as important. Sovereignty, for us, combines two dimensions: how we built the model, and how it transfers to our customers. We offer full supply-chain integrity and account for every decision, from data ingestion, through pre- and post-training, to the final evaluations. We provide transparency. Customers have full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property of the model. Read our tech report for full details. What Kolibri Delivers We optimized Kolibri for performance across a wide range of sectors considering their particular domain-specific language, regulatory, and procedural realities. Its small and efficient size provides our customers with flexibility to run it efficiently on-premise, without sending internal data to third-party inference services. The spotlight in this section introduces the model's capabilities, before we describe them in section How we built Kolibri at high velocity. Foundational capabilities for enterprise and government With Kolibri we optimize the trade-off between model capability and deployment costs, using 3B active parameters out of 78B total. Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German. The Pareto frontier is a concept from economics, marking the best achievable combinations of two objectives, where improving on one means giving up some of the other. None of the compared models delivers more quality at the same serving cost, or the same quality at lower cost. Average benchmark score [%]EnglishGermanKolibriKolibri OriginOther post-trained modelsPareto frontier Performance vs. throughput for post-trained models in English (left) and German (right). Metrics show unweighted average benchmark scores against decoded text per second and GPU. Higher and further right is better. Across math, coding, grounding, and long-context tasks, Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super. AIME 2025MathAIME 2025 (DE)MathAIME 2026MathAIME 2026 (DE)MathGPQA (diamond)KnowledgeGPQA (diamond, DE)KnowledgeAA-Omniscience IndexGrounding / hallucinationspublic set; from −100 to 100BrowseCompAgenticτ³-bench bankingAgenticτ²-bench retailAgenticτ²-bench airlineAgenticτ²-bench telecomAgenticBFCL v4 overallAgenticLiveCodeBench v6CodeHumanEval+CodeLongBench ProLong contextAA-LCRLong contextShow the numbersBenchmarkKolibriKolibri OriginQwen3.6-35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6BAIME 202596.981.984.691.779.8AIME 2025 (DE)87.573.582.985.672.3AIME 202696.081.591.090.483.1AIME 2026 (DE)90.075.284.487.578.5GPQA (diamond)84.368.183.478.074.7GPQA (diamond, DE)81.358.580.676.672.9AA-Omniscience Index-32.8-64.0-15.3-36.5-24.0BrowseComp29.44.426.929.1–τ³-bench banking38.15.710.615.55.7τ²-bench retail69.958.571.667.562.9τ²-bench airline76.758.770.772.740.0τ²-bench telecom94.767.599.168.141.5BFCL v4 overall61.436.467.261.058.0LiveCodeBench v685.959.282.582.071.2HumanEval+92.776.892.894.792.8LongBench Pro64.5–70.862.956.4AA-LCR68.3–69.767.052.3 Foundational capabilities of Kolibri. Kolibri is a balanced generalist model with competitive performance across math, code, long context, agentic capabilities, and knowledge. Benchmark scores are shown on a shared 0–100 scale, higher is better. Contextualized performance for real-world applications Public benchmarks fail to capture specialized sector needs, so we developed our own internal evaluation suites for the verticals that matter to our customers, such as the German public sector, aviation, manufacturing and the automotive industry. Each suite mirrors the skills, workflows, and edge cases required in these sectors, and paired synthetic training environments let us improve Kolibri against these evaluations without ever training on customer data. Read more in section Contextualized performance. Score on internal customer-proxy benchmarkAutomotive supplier 0.72 → 0.99Semiconductors 0.35 → 0.80German public sector 0.54 → 0.75Industrial drive technology 0.31 → 0.60Aerospace 0.14 → 0.59one checkpoint, one evalmean of that dayKolibri OriginKolibri Contextualized performance across successive post-training runs on internal customer-proxy benchmarks. Dots are single evaluations of training checkpoints, lines are the mean of each day's evaluations. Performance climbs across all five verticals, higher is better. Answers that are grounded in your documents. We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context. Our customers value and request this feature, so we continuously track and validate abstention accuracy. We provide more details in the Grounding section. Native German and English model. We developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German. We used translation sparingly (6% overall), since translated text tends to carry the cultural fingerprint of its source language. The result is a model that is bilingual by design, not an English model that has read some German. Read more in sections German pre-training data and Specialized tokenizers. Control and compliance by design: upholding our customers' sovereignty We built Kolibri with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up, with copyright law being a focus of our work to trustworthy technology. We are transparent about our model weights and about curation of our training data, so that decisions behind development are visible. Through its reasoning traces, it becomes explainable how the model came to a particular answer. With our Merlin-Arthur protocol the model grounds trustworthiness: Kolibri refrains from answering when the context doesn't support an answer. Our teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control. We own the entire pipeline, from data curation, through pre- and post-training, to optimization in our Model Factory. This control of our end to end pipeline underpins the model's sovereignty. Control also passes on to our customers and Kolibri's small size gives them full freedom of deployment and supports controllable reasoning effort to trade cost and latency against answer quality. How We Built Kolibri at High Velocity Our Model Factory: two models, three months apart Our Model Factory is our answer to iteration speed, minimizing the time it takes to go from knowing what a model gets wrong to training one that does better. We implemented the training pipeline as code, so that the learnings of our team landed in one versioned training recipe instead of scattered scripts and notes. Here is what that looked like between Kolibri Origin and Kolibri. Work on the pipeline began in January. Five months and hundreds of ablation runs later, Kolibri Origin finished pre-training at target scale on 11 June. Kolibri finished on 11 September. In the three months between those two dates we went from 30B parameters to 78B, from a 65k-token context window to up to 1M, and from 7.5T training tokens to 20T. To get those 20T, the pipeline processed over 200T tokens of raw data, filtering, deduplicating, and curating it down to what we actually trained on. We changed the attention design, tripled the number of experts, increased sparsity, replaced the routing algorithm, improved our post-training data, more than doubled the number of environment tasks, and taught the model to reason at four different effort levels. Both models kept a similar number of parameters active per token, yet Kolibri trained faster per token than Kolibri Origin thanks to the work of our efficiency team.