Hacker News (100+) - 03 Oct 2026

Page 2 of 2

In the end German entered Kolibri as a 2.4T-token unique pool, 80% of it curated or generated by us and 20% from open datasets, and the model saw it at 21.3% of pre-training tokens, roughly 4.3T over the 20T run through upsampling. Each German token was seen 1.8 times on average, well inside the four-epoch limit past which repetition stops paying off. Around 85% of the German the model read is web text, either organic or rephrased from organic. The rest is curated documents - parliamentary proceedings, legal texts and other data in the public domain - and that small translated share.

Specialized tokenizers for English-German

We built a bilingual English-German tokenizer, trained and specialized on each model's pre-training dataset. Because of the 21% German share in our data, it compresses German language better than other SOTA models. This led to more efficient inference (fewer tokens) minimizing costs and shortening response times. We introduce a new way to train tokenizers, UniBPE, that respects the morphology of languages better than existing approaches, especially the compound structure of German, without sacrificing English token efficiency.

The two leading approaches to train tokenizers are BPE and Unigram. We combine those two, keeping the bottom-up approach of BPE and using the Unigram training objective for selecting which merge to add to the vocabulary. On a 128k vocabulary trained on our English/German dataset, this substantially improves tokenization.

Grounding: reducing hallucinations

Common LLM training and evaluation rewards guessing: a guess has some chance of landing the correct answer, while abstention has none. Therefore, models learn to answer with whatever they have, even if their input does not provide enough information to warrant a response. Asking a model to provide citations often backfires for the same reason: models can simply hallucinate plausible-looking sources to justify an ungrounded answer.

For a regulated customer, a model that knows to abstain is the difference between a pilot and a deployment. We treat saying "I don't know" as an important model capability and (1) develop and track dedicated grounding and anti-hallucination measures and (2) use and develop dedicated training procedures to improve this abstention capability.

During training, we use training data samples where the correct answer is "I don't know". While fine-tuning for abstention is gaining traction industry-wide, high-quality negative examples remain hard to come by, especially where our customers need them, in narrow verticals where all data is scarce. Therefore, we also train with our Merlin-Arthur procedure, developed in-house and explained in detail in our blog post, which solves data scarcity by automatically exploiting the model's weaknesses at each training step and generates synthetic negative examples from available documents. This training becomes a game with three players. Arthur is the model we train and ship. Merlin takes an existing document context, generates a new one that increases Arthur's probability of answering correctly. Morgana generates a new datapoint by stripping out the relevant evidence from the document and tries to lure Arthur into a hallucination. Arthur does not know which one he's facing, so his valid strategy becomes to carefully read the question and the context, to be able to answer on Merlin's context, while recognizing that abstention is the strictly required answer for Morgana's redacted context - making any guess, even a lucky one, incorrect.

For measuring hallucination abstention, we use established benchmarks alongside our own proxies derived from customer use cases. We've also developed our own "M/A grounding score" which falls out of our Merlin-Arthur setup representing a lower bound on how much of the answer provably came from the document.

Kolibri hallucinates far less than Kolibri Origin: it abstains instead of answering wrong on 44% of AA-Omniscience items (Origin: 15%), and on RGB it holds back more often (86% vs 74%) and invents fewer falsehoods (87% vs 76%). On our own M/A grounding score, which certifies rather than estimates how much of an answer came from the document, Kolibri reaches 0.23 where Kolibri Origin, along some other models, reach 0.

Show the numbers

| Benchmark | Kolibri | Kolibri Origin | Qwen3.6-35B-A3B | Qwen3-Next 80B-A3B | Nemotron 3 Super 120B-A12B | Mistral Small 4 119B-A6B |

|---|---|---|---|---|---|---|

| AA-Omniscience Non-Hallucination Rate | 44.0 | 14.8 | 56.7 | 12.3 | 13.9 | 34.7 |

| RGB: holds back | 85.6 | 73.9 | 79.6 | 81.3 | 74.6 | 82.3 |

| RGB: invents nothing | 87.3 | 75.6 | 84.3 | 83.9 | 86.0 | 87.0 |

| FRAMES | 71.2 | 65.7 | 74.7 | 68.9 | 74.9 | 71.9 |

| M/A grounding score | 0.23 | 0.00 | 0.12 | 0.00 | 0.00 | 0.06 |

Contextualized performance

Most public benchmarks fail to capture specialized sector needs. To measure contextualized performance - meaning how well a model handles the unique workflows and domain knowledge of specific industries - we need measurements of performance in specialized verticals which are important to our customers, such as the German public sector, legal, hardware, consumer electronics, and automotive.

For this, we first built domain-specific evaluation suites that mirror the skills, tool-calls, workflows, and edge cases required in key sectors and thus capture contextualized performance. Once this was done, we created training environments by generating high-quality specialized synthetic data capturing those required skills. To increase the real-world robustness of the model on those tasks, we additionally randomized the environments around concrete details (tools, configurations, harnesses).

Together, these contributed to an iterative engine that allowed us to sharpen the model's agentic RAG capabilities, domain-specific logic and capabilities, and to hill-climb the contextualized evaluations without ever training on customer data.

These evaluation suites run automatically as part of our pipeline: every checkpoint is scored on them as it lands. Kolibri compares well against competing open-weight models. On the agentic-RAG benchmark Honeypot, Kolibri outperforms all compared models, including the much larger Nemotron 3 Super and Mistral Small 4. On a cleaned version of the agentic-RAG benchmark MuSiQue, it is second only to the more massive Nemotron 3 Super and well ahead of the rest. On the five customer applications, Kolibri leads or is on par with the best model in four of five. These findings speak to our ability to handle a range of data and harness setups, across use-cases, whereas the compared models are sensitive to these details.

These evaluations for measuring contextualized performance can exist because customers tell us how their applications fall short. Each one encodes what we learned from those conversations: what kind of documents matter, which tools the model should get, which questions are hard, where the deployed system frustrates its users today. That makes the exchange concrete in both directions. A customer who shows us a failure case gets it turned into a benchmark that every future checkpoint is measured against, and we can hill-climb the target without ever training on customer data or risking over-fitting.

Show the numbers

| Benchmark | Kolibri | Kolibri Origin | Qwen3-Next 80B-A3B | Qwen3.6-35B-A3B | Nemotron 3 Super 120B-A12B | Mistral Small 4 119B-A6B |

|---|---|---|---|---|---|---|

| MuSiQue (cleaned) | 77.3 | 42.7 | 50.5 | 61.2 | 79.1 | 66.8 |

| Honeypot | 80.8 | 25.3 | 13.5 | 74.3 | 68.8 | 68.1 |

| Semiconductors | 80.4 | 35.3 | 41.2 | 79.4 | 69.6 | 62.7 |

| German public sector | 75.0 | 54.0 | 29.5 | 72.0 | 78.0 | 50.0 |

| Aerospace | 58.9 | 14.1 | 48.1 | 59.0 | 54.9 | 47.0 |

| Automotive supplier | 99.0 | 72.4 | 84.2 | 92.6 | 91.0 | 87.1 |

| Industrial drive technology | 60.0 | 31.4 | 32.7 | 59.5 | 37.3 | 56.8 |

Benchmarks

Benchmarks were run using our own harnesses and, where applicable, all models used the highest respective reasoning effort.

| Type | MoE | | | | | | | MoE | | | MoE | | Dense | |

|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|

| Active parameters | 3B | | | | | | | 4-6B | | | 12B | | 27B * 70B | |

| Benchmark | Kolibri | Kolibri Origin | GLM-4.7 Flash 30B-A3B | Nemotron 3 Nano 30B-A3B | Qwen3.5 35B-A3B | Qwen3.6 35B-A3B | Qwen3-Next 80B-A3B Thinking | Gemma 4 26B-A4B IT | GPT-OSS 120B | Mistral Small 4 119B-A6B | GLM-4.5 Air 106B-A12B | Nemotron 3 Super 120B-A12B | Qwen3.8 27B | Apertus 70B Instruct |

| Overall (EN) | 75.5 | 54.1 | 64.7 | 65.6 | 74.7 | 71.4 | 62.4 | 71.9 | 72.3 | 63.1 | 64.4 | 73.0 | 80.2 | - |

| Overall (DE) | 70.8 | 46.4 | 50.4 | 59.3 | 69.8 | 67.3 | 58.0 | 66.3 | 70.2 | 61.4 | 64.8 | 67.9 | 79.9 | - |

| Knowledge | | | | | | | | | | | | | | |

| Average (EN) | 50.1 | 39.7 | 45.5 | 46.0 | 52.7 | 52.1 | 48.4 | 51.4 | 50.0 | 47.5 | 45.7 | 52.0 | 56.8 | - |

| Average (DE) | 57.6 | 44.9 | 46.5 | 41.2 | 61.3 | 61.0 | 55.7 | 61.5 | 58.0 | 51.4 | 52.2 | 59.5 | 69.2 | - |

| GPQA Diamond (EN) | 84.3 | 68.1 | 73.1 | 73.9 | 83.8 | 83.4 | 76.1 | 81.1 | 76.4 | 74.7 | 73.2 | 78.0 | 89.2 | 29.5 |

| GPQA Diamond (DE) | 81.3 | 58.5 | 59.8 | 49.6 | 84.2 | 80.6 | 72.2 | 80.1 | 76.0 | 72.9 | 71.1 | 76.6 | 88.1 | 31.4 |

| Humanity's Last Exam (EN) | 21.5 | 9.4 | 15.4 | 12.1 | 20.4 | 21.1 | 11.6 | 19.2 | 19.4 | 9.7 | 8.7 | 20.6 | 35.6 | 5.2 |

| Humanity's Last Exam (DE) | 15.9 | 10.4 | 9.1 | 13.1 | 18.1 | 20.5 | 15.6 | 23.4 | 20.7 | 10.5 | 10.5 | 22.3 | 37.2 | 5.7 |

| AA-Omniscience Accuracy (public set) | 14.8 | 11.3 | 17.0 | 19.5 | 22.0 | 19.5 | 24.2 | 20.7 | 23.3 | 25.0 | 20.0 | 26.7 | 17.5 | 13.5 |

| AA-Omniscience Index (public set) | -32.8 | -64.0 | -62.8 | -45.7 | -47.3 | -15.3 | -42.3 | -47.3 | -35.2 | -24.0 | -28.8 | -36.5 | -9.5 | - |

| MMLU-Pro CoT (EN) | 80.0 | 70.1 | 76.5 | 78.3 | 84.6 | 84.3 | 81.7 | 84.5 | 80.8 | 80.4 | 80.9 | 82.7 | 85.0 | 43.0 |

| MMLU-ProX CoT (DE) | 75.5 | 65.7 | 70.7 | 61.0 | 81.7 | 81.9 | 79.4 | 81.1 | 77.2 | 70.7 | 74.9 | 79.7 | 82.4 | 37.3 |

| Math | | | | | | | | | | | | | | |

| Average (EN) | 96.5 | 81.7 | 88.8 | 88.8 | 90.1 | 87.8 | 86.3 | 87.4 | 90.7 | 81.4 | 82.8 | 91.1 | 97.8 | - |

| Average (DE) | 88.8 | 74.3 | 45.2 | 84.3 | 79.6 | 83.7 | 83.7 | 88.1 | 90.8 | 75.4 | 80.6 | 86.5 | 96.7 | - |

| AIME 2025 (EN) | 96.9 | 81.9 | 89.4 | 89.6 | 88.1 | 84.6 | 84.2 | 87.3 | 90.6 | 79.8 | 81.9 | 91.7 | 97.9 | 0.6 |

| AIME 2025 (DE) | 87.5 | 73.5 | 43.8 | 84.4 | 76.7 | 82.9 | 80.6 | 88.1 | 90.6 | 72.3 | 80.6 | 85.6 | 96.5 | 0.2 |

| AIME 2026 (EN) | 96.0 | 81.5 | 88.3 | 87.9 | 92.1 | 91.0 | 88.5 | 87.5 | 90.8 | 83.1 | 83.8 | 90.4 | 97.7 | 0.6 |

| AIME 2026 (DE) | 90.0 | 75.2 | 46.7 | 84.2 | 82.5 | 84.4 | 86.7 | 88.1 | 91.0 | 78.5 | 80.6 | 87.5 | 96.9 | 0.0 |

| Agentic | | | | | | | | | | | | | | |

| Average (EN) | 63.4 | 41.6 | 58.9 | 46.4 | 63.4 | 62.1 | 46.3 | 54.6 | 54.0 | 40.7 | 53.5 | 54.9 | 66.7 | - |

| TerminalBench 2.1 | 27.7 | - | 20.2 | 9.7 | 39.7 | - | 8.6 | - | 29.2 | 21.0 | - | 39.7 | 76.8 | - |

| Tau2-Bench (Telecom) | 94.7 | 67.5 | 95.9 | 45.9 | 97.7 | 99.1 | 43.9 | 45.3 | 73.1 | 41.5 | 53.8 | 68.1 | 82.5 | 10.8 |

| Tau2-Bench (Retail) | 69.9 | 58.5 | 57.9 | 64.9 | 70.8 | 71.6 | 60.8 | 71.3 | 60.5 | 62.9 | 61.4 | 67.5 | 68.7 | 9.6 |

| Tau2-Bench (Airline) | 76.7 | 58.7 | 68.7 | 52.7 | 76.0 | 70.7 | 65.3 | 73.3 | 72.7 | 40.0 | 70.7 | 72.7 | 83.3 | 40.0 |

| Tau3-Bench (Banking) | 38.1 | 5.7 | 7.2 | 5.7 | 11.3 | 10.6 | 5.4 | 16.0 | 14.7 | 5.7 | 6.4 | 15.5 | 50.0 | 2.1 |

| BFCL v3 (multi-turn) | 39.8 | 22.8 | 58.2 | 47.9 | 54.0 | 53.5 | 51.4 | 53.4 | 45.6 | 36.2 | 61.6 | 44.6 | 42.5 | 0.6 |

| BFCL v4 (overall) | 61.4 | 36.4 | 65.4 | 61.5 | 70.5 | 67.2 | 51.0 | 68.2 | 57.3 | 58.0 | 67.2 | 61.0 | 73.2 | - |

| BFCL v4 (non-live AST) | 79.1 | 78.1 | 83.3 | 85.0 | 85.8 | 88.2 | 83.6 | 83.7 | 35.8 | 83.6 | 85.5 | 45.0 | 85.3 | - |

| BFCL v4 (live) | 78.9 | 73.7 | 78.3 | 78.8 | 80.2 | 81.4 | 82.5 | 80.2 | 70.4 | 78.4 | 78.2 | 77.6 | 79.9 | - |

| BFCL v4 (multi-turn) | 47.5 | 27.5 | 62.7 | 53.5 | 59.9 | 58.1 | 56.0 | 61.4 | 55.4 | 40.4 | 65.2 | 51.7 | 55.5 | - |

| BFCL v4 (memory) | 62.8 | 19.4 | 41.5 | 39.1 | 62.6 | 53.8 | 35.3 | 52.9 | 50.7 | 39.1 | 43.4 | 59.6 | 79.6 | - |

| BFCL v4 (web search) | 62.5 | 10.5 | 69.0 | 66.0 | 75.0 | 68.5 | 12.5 | 75.0 | 57.0 | 69.0 | 71.0 | 71.5 | 82.0 | - |

| BrowseComp | 29.4 | 4.4 | - | 14.5 | 36.5 | 26.9 | 2.8 | 25.5 | 31.2 | - | - | 29.1 | 46.4 | - |

| Code | | | | | | | | | | | | | | |

| Average (EN) | 89.3 | 68.0 | 67.8 | 81.8 | 85.0 | 87.7 | 83.6 | 89.0 | 90.8 | 82.0 | 79.8 | 88.3 | 94.2 | - |

| LiveCodeBench v6 | 85.9 | 59.2 | 46.5 | 71.3 | 77.8 | 82.5 | 73.9 | 82.3 | 87.5 | 71.2 | 67.8 | 82.0 | 93.8 | 8.7 |

| HumanEval+ | 92.7 | 76.8 | 89.0 | 92.4 | 92.2 | 92.8 | 93.3 | 95.7 | 94.1 | 92.8 | 91.8 | 94.7 | 94.7 | 41.6 |

| SWE-Bench Verified | 66.4 | - | 51.0 | 38.6 | 71.6 | 73.8 | - | 57.8 | - | 60.8 | 11.6 | 60.2 | 72.6 | - |

| Instruction Following | | | | | | | | | | | | | | |

| Average (EN) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | - |

| IFBench (loose-prompt) | 78.1 | 62.5 | 64.5 | 73.2 | 72.7 | 66.1 | 60.7 | 79.9 | 71.1 | 49.8 | 38.2 | 73.7 | 81.9 | 25.5 |

Get Started

Our model is available openly on Hugging Face under an Apache 2.0 License.

Kolibri requires the aleph-alpha-inference package that provides the Kolibri

vLLM plugin. You can either use the provided container image ghcr.io/aleph-alpha/aleph-alpha-inference, or install the package from Aleph-Alpha/aleph-alpha-inference, which also installs the vLLM version it supports:

pip install "aleph-alpha-inference>=1.0"

Serve the model with reasoning and tool-calling enabled.

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \

--reasoning-parser kolibri1 \

--tool-call-parser kolibri1 \

--enable-auto-tool-choice

To serve contexts beyond 262,144 tokens, add --max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'. The recommended sampling parameters for the model are temperature=1.0, top_p=0.97 and top_k=128.

Contact for Deployment and Specialization

Contact our team who will be happy to support you through our enterprise deployment and specialization options: contact sales.

Previous page | HN comments | More Hacker News (100+) | Headlines

Original: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/