Hacker News (100+) - 03 Oct 2026
Page 1 of 2
Aleph Alpha
Kolibri Has Landed: A Sovereign Open-Weight Model
On the Day of German Reunification, we are releasing our new model: Kolibri.
Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active. It supports context lengths of up to 1M tokens. The model can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms.
Kolibri is the result of continuous iteration of our model training effort. We first built a model training pipeline and validated it by building Kolibri Origin, a 30B total, 3B active model with a much shorter 65k token context window. Kolibri ran through the same pipeline: from data ingestion and curation, through ablations, pre-training, and post-training, to the final evals. It enabled running hundreds of ablation experiments and stable pre-training that ran without a person having to step in when hardware failed or a data connection dropped. We continuously monitored training metrics and standardized monitoring for custom benchmarks. The time we put into building and iterating on this pipeline was a valuable investment. We see it in how much better Kolibri is than Kolibri Origin, and in how little time separates their releases.
Kolibri is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace. We specialized Kolibri for German, reasoning, math, agentic behavior, and further capabilities our customers need in production. The aim of this specialization was to optimize performance in our customers' specific use cases. Through specialization, customers achieve contextualized performance in their AI operations and they can monitor its economic impact, so that ROI stays measurable and grows over time.
Specialization alone is not enough. Sovereignty is just as important. Sovereignty, for us, combines two dimensions: how we built the model, and how it transfers to our customers. We offer full supply-chain integrity and account for every decision, from data ingestion, through pre- and post-training, to the final evaluations. We provide transparency. Customers have full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property of the model.
Read our tech report for full details.
What Kolibri Delivers
We optimized Kolibri for performance across a wide range of sectors considering their particular domain-specific language, regulatory, and procedural realities. Its small and efficient size provides our customers with flexibility to run it efficiently on-premise, without sending internal data to third-party inference services. The spotlight in this section introduces the model's capabilities, before we describe them in section How we built Kolibri at high velocity.
Foundational capabilities for enterprise and government
With Kolibri we optimize the trade-off between model capability and deployment costs, using 3B active parameters out of 78B total. Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German. The Pareto frontier is a concept from economics, marking the best achievable combinations of two objectives, where improving on one means giving up some of the other. None of the compared models delivers more quality at the same serving cost, or the same quality at lower cost.
Across math, coding, grounding, and long-context tasks, Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super.
Show the numbers
| Benchmark | Kolibri | Kolibri Origin | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B | Mistral Small 4 119B-A6B |
|---|---|---|---|---|---|
| AIME 2025 | 96.9 | 81.9 | 84.6 | 91.7 | 79.8 |
| AIME 2025 (DE) | 87.5 | 73.5 | 82.9 | 85.6 | 72.3 |
| AIME 2026 | 96.0 | 81.5 | 91.0 | 90.4 | 83.1 |
| AIME 2026 (DE) | 90.0 | 75.2 | 84.4 | 87.5 | 78.5 |
| GPQA (diamond) | 84.3 | 68.1 | 83.4 | 78.0 | 74.7 |
| GPQA (diamond, DE) | 81.3 | 58.5 | 80.6 | 76.6 | 72.9 |
| AA-Omniscience Index | -32.8 | -64.0 | -15.3 | -36.5 | -24.0 |
| BrowseComp | 29.4 | 4.4 | 26.9 | 29.1 | - |
| 3-bench banking | 38.1 | 5.7 | 10.6 | 15.5 | 5.7 |
| 2-bench retail | 69.9 | 58.5 | 71.6 | 67.5 | 62.9 |
| 2-bench airline | 76.7 | 58.7 | 70.7 | 72.7 | 40.0 |
| 2-bench telecom | 94.7 | 67.5 | 99.1 | 68.1 | 41.5 |
| BFCL v4 overall | 61.4 | 36.4 | 67.2 | 61.0 | 58.0 |
| LiveCodeBench v6 | 85.9 | 59.2 | 82.5 | 82.0 | 71.2 |
| HumanEval+ | 92.7 | 76.8 | 92.8 | 94.7 | 92.8 |
| LongBench Pro | 64.5 | - | 70.8 | 62.9 | 56.4 |
| AA-LCR | 68.3 | - | 69.7 | 67.0 | 52.3 |
Contextualized performance for real-world applications
Public benchmarks fail to capture specialized sector needs, so we developed our own internal evaluation suites for the verticals that matter to our customers, such as the German public sector, aviation, manufacturing and the automotive industry. Each suite mirrors the skills, workflows, and edge cases required in these sectors, and paired synthetic training environments let us improve Kolibri against these evaluations without ever training on customer data. Read more in section Contextualized performance.
Answers that are grounded in your documents. We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context. Our customers value and request this feature, so we continuously track and validate abstention accuracy. We provide more details in the Grounding section.
Native German and English model. We developed a bilingual German/English tokenizer and focused on including organic German data throughout the training process of the model, so that 21.3% of the pre-training tokens are German. We used translation sparingly (6% overall), since translated text tends to carry the cultural fingerprint of its source language. The result is a model that is bilingual by design, not an English model that has read some German. Read more in sections German pre-training data and Specialized tokenizers.
Control and compliance by design: upholding our customers' sovereignty
We built Kolibri with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up, with copyright law being a focus of our work to trustworthy technology.
We are transparent about our model weights and about curation of our training data, so that decisions behind development are visible. Through its reasoning traces, it becomes explainable how the model came to a particular answer. With our Merlin-Arthur protocol the model grounds trustworthiness: Kolibri refrains from answering when the context doesn't support an answer.
Our teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control. We own the entire pipeline, from data curation, through pre- and post-training, to optimization in our Model Factory. This control of our end to end pipeline underpins the model's sovereignty. Control also passes on to our customers and Kolibri's small size gives them full freedom of deployment and supports controllable reasoning effort to trade cost and latency against answer quality.
How We Built Kolibri at High Velocity
Our Model Factory: two models, three months apart
Our Model Factory is our answer to iteration speed, minimizing the time it takes to go from knowing what a model gets wrong to training one that does better. We implemented the training pipeline as code, so that the learnings of our team landed in one versioned training recipe instead of scattered scripts and notes. Here is what that looked like between Kolibri Origin and Kolibri.
Work on the pipeline began in January. Five months and hundreds of ablation runs later, Kolibri Origin finished pre-training at target scale on 11 June. Kolibri finished on 11 September. In the three months between those two dates we went from 30B parameters to 78B, from a 65k-token context window to up to 1M, and from 7.5T training tokens to 20T. To get those 20T, the pipeline processed over 200T tokens of raw data, filtering, deduplicating, and curating it down to what we actually trained on. We changed the attention design, tripled the number of experts, increased sparsity, replaced the routing algorithm, improved our post-training data, more than doubled the number of environment tasks, and taught the model to reason at four different effort levels. Both models kept a similar number of parameters active per token, yet Kolibri trained faster per token than Kolibri Origin thanks to the work of our efficiency team.
| | Kolibri Origin | Kolibri |
|---|---|---|
| Finished pre-training | 11 June 2026 | 11 September 2026 |
| Release | no public release | 3 October 2026 |
| Reasoning mode | Yes (one mode only) | Yes (none, low, medium, high) |
| Total parameters | 30.6B | 78.1B |
| Active parameters / token | 3.27B | 3.46B |
| Pre-training tokens | 7.51T | 20T |
| Layers | 50 (2 dense + 48 MoE, 1 shared expert) | 50 (all MoE, 1 shared expert) |
| Pre-training context length | 8,192 (8k) | 16,384 (16k) |
| Longest trained length | 65,536 (64k) | 262,144 (256k) |
| Tokenizer vocabulary | 96,000 | 128,000 |
| Model dimension | 2,048 | 2,560 |
| Attention heads (query / KV) | 32 / 4 | 48 / 4 |
| Experts (total / active) | 128 / 8 | 384 / 6 |
| Expert hidden dim | 768 | 512 |
| Attention pattern | full attention, all layers | sliding window (512) + full attention every 5th layer |
| Knowledge cutoff | EN: 1 Sept 2024, DE: 1 Aug 2025 | EN/DE: 18 Jun 2026 |
What made this possible is our training pipeline, an effort to convert research-grade model development into a fully automated production-ready infrastructure to design and train large language models. Our pipeline codebase was shared by all. Every proposed code change triggers a small end-to-end model run; training, evaluation, to find out if something broke within minutes. Our runs are GitHub Actions workflows, and reproducing one means checking out a commit. We used the pipeline for the main training and the hundreds of ablations that ran to make our architectures and data-mix decisions.
Training checkpoints land roughly every hour and the pipeline automatically evaluates them on English and German knowledge, maths, code, instruction following, tool use, long context, safety, abstention to hallucination and grounding. We can watch capabilities appear every day rather than finding out how the model turned out at the end. The run itself held up better than we expected. Over 21 days of pre-training of Kolibri, we hit 38 unplanned interruptions, roughly one per 10,000 GPU-hours, caused by hardware faults or a connection timing out. Those were automatically handled by the pipeline without manual intervention. The cluster automatically restarted the training job(s) on a different set of nodes, and training picked up from a checkpoint at most 250 steps back.
Having access to the whole pipeline, with checkpoints available to everyone, means that any team can own a capability end to end rather than a stage of an assembly line. One team trained and evaluated the grounding capability of Kolibri as a single piece of work, seamlessly integrating into the overall model.
However, getting here was not a linear walk. We stumbled. We stopped the Kolibri Origin pre-training after a few trillion tokens and restarted it from scratch, as we identified a data-shuffling bug that escaped our tests. We ran ablations, took decisions, only to find a bug or an error in the configuration afterwards, forcing us to rerun some experiments. Alongside our own experience, we tracked state-of-the-art architectures, best practices, and the latest research advances in LLMs. Those accumulated learnings led to improvements in our processes and guardrails in our pipeline.
Looking at the delta between Kolibri Origin and Kolibri, the improvement is notable, but this is also the easy direction: more parameters, more data, a well-understood architecture family, and still at small scale. As we are contemplating scaling up, we are happy to face challenges on our foundation. But the most durable thing we built this year is not the pipeline, it is a team with the proven capability to build, post-train, and ship LLMs from raw data at high velocity.
Architecture and pre-training
Kolibri has 78B total parameters, which is 2.5 times more than Kolibri Origin, with ~3B active parameters. In our experiments, increasing the size of our model from 32B to 123B led to ever-improved performance. Yet the larger size came with larger training and serving costs. The latter drove the decision: 123B can handle only 3 long-context 256k-token user queries on two H100s, while 78B handles 18 concurrent requests and decodes 28% faster. We used 384 smaller experts rather than fewer wide ones, as they performed better in our tests. We applied the same efficiency-first reasoning to attention. Out of 50 total layers, only 10 process full context, while the remaining 40 use a tight 512-token focused window. This keeps decode computation and memory bounded in those layers regardless of context length. The efficiency gains benefit both serving the trained model and training during post-training RL, which relies on huge amounts of inference.
We trained Kolibri on 768 B200 GPUs in three stages: 20T tokens of pre-training at a 16k sequence length over 21 days, 3.44T tokens of mid-training at 64k, and 200B tokens of long-context adaptation at 256k. This is nearly 24T tokens in total, roughly three times what Kolibri Origin consumed. German accounts for more than a fifth of the pre-training mix, about 4.3T tokens, against roughly 62% English and 14% code. Compared to pre-training, mid-training data is a much more selective pool of curated datasets weighted toward reasoning, problem-solving, code and agentic data. For long-context, rather than train on long documents alone, which tends to erode the skills acquired earlier, we interleaved long documents with the high-quality mid-training data of the previous stage. For the long-context mix, we removed synthetic long documents to avoid artificially inflating benchmarks such as RULER.
We optimized with Muon, as we did for Kolibri Origin. We put special focus on training stability, resulting in a robust training run without any loss spikes for either model. With Kolibri, for routing between experts during training, we introduce exact quantile balancing. Quantile balancing was introduced in Kimi K3, where the global quantile is estimated from histograms because an exact computation was considered too expensive to communicate; we show that it can be computed exactly at fixed cost independent of batch size, and that the exactness improves both load balance and model quality.
Post-training at scale
Post-training happened in two stages; the first step is supervised fine-tuning to teach the model core reasoning ability and how to interact in a chat, followed by large-scale reinforcement learning to train the model to reason across a diverse suite of long-horizon tasks. We built a pipeline for both of these steps, which allowed us to exercise fine-grained control over the behavior of the model.
For SFT we generated a total of 174B tokens worth of synthetic data, which we filtered for quality and combined with filtered versions of permissively licensed open-source datasets to obtain a high-quality training mix of 268B tokens. For reinforcement learning we trained on a broad set of environments that contained more than 1.2 million curated tasks across diverse domains such as code & math reasoning, agentic tasks, instruction following, question answering, tool calling, and more. We optimized our in-house training codebase for high-performance and use asynchronous training, where we generate training data on our environments using the current model in parallel to training the model on already generated data.
Across both stages, we taught the model to reason at different effort levels (none, low, medium, high), which means that the user can exercise control over how much compute the model should invest to find a solution to the task at hand. This allows our customers to trade-off cost and inference speed against the quality of the final answer.
German pre-training data
From our experiments on small proxy models, we found that training with around 20% German data leads to optimal results, which meant we needed to find 4T German tokens to train our model at a 20T horizon. Open German datasets help, but are far from enough: after deduplication and filtering, we were left with 390B German tokens, well short of our goal. As we argued in Sauerkraut, Not Burgers, German capability has to come from high-quality German texts; relying heavily on machine-translations would lead to poor results due to subtle translation errors and a lack of authentic German cultural context. We closed the token gap in three different ways.
The first was to curate German from Common Crawl ourselves. We built a pipeline specialized for German data. German is not English, and German data cannot be filtered like English data: we had to retune filtering parameters for the German language. One typical filter in a language data pipeline is to remove documents with too many long words, but German administrative prose routinely exceeds the English bound on mean word length, so the standard settings quietly remove the register that public administration writes in. After retuning, our German pipeline gave us 1.3T unique tokens of organic German web.
The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content. This teaches the model the same facts in several surface forms and multiplies the information contained in scarce data, and it is a different operation from translation: the source is German, so the subject matter and the cultural affinity stay German: chancellor, not president. It does not add much new knowledge, rather new phrasings of knowledge that was already in the corpus. Rephrasing gave us about 1T unique tokens, making it the single largest source of German in the model.
The third was translation, and this we only used in Kolibri Origin. Translating English into German works when the model, the prompts and the chunking are chosen carefully, but it carries two problems. Output can still show translationese, the literal rendering of idioms: "Drive safe!" becomes "Fahre sicher!" rather than "Komm gut an!". More importantly, cultural context does not translate. A corpus translated from English inherits the geographic, demographic and institutional distribution of the English web, so a model trained on it speaks German about a world that looks American.
Next page | HN comments | More Hacker News (100+) | Headlines
Original: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/