Thank you Aleph Alpha team for making it open.
We as many other’s were curious to try and benchmark it.
On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.
No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/
>We trained Kolibri with abstention data and with our Merlin-Arthur protocol. As a result, it is trained to say "I don't know" when the answer isn't in the context.
https://aleph-alpha.com/en/blog/bounding-hallucinations-merl...
The thing to note here, besides the transparency and the fact that it’s actually a good model that also works well on coding and agentic tasks, is that it’s the first release by a team formed less than a year ago, with a strong focus on iteration velocity. There’s more to come.
disclaimer: I‘m part of the training team, happy to answer any questions
Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark.
Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?
For a post to make such a big deal about sovereignty it is a bit misleading to not mention that the company is slated to be merged with Cohere, a Canadian company.
And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.
Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.
I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models.
Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.
So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.
If the second model is cheap and fast enough, there is a business model.
You don’t even need to audit all the intermediate steps, just tool calls and end results.
The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.
I get that doesn't invalidate the real "point" of the model, but...
> 3B active
> built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace
Cool to be doing more independent model lineages, but I sure hope no one actually uses this as part of any aerospace engineering...
I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
Related ongoing thread:
Aleph Alpha Kolibri: How the sovereign German LLM works - https://news.ycombinator.com/item?id=49943034 - Oct 2026 (162 comments)
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.
This thing is worse than a Qwen3.8 27B.
"Languages: German and English" – this is odd. That means their dataset is limited. In my understanding, frontier models are trained on multilingual datasets and can combine knowledge no matter what language it was written in.
Looks interesting. I just went to download from HF, but they only have fp16 which won't fit on my Mac.
Good to see Europe adding toe what Mistral is doing. +100
78B MoE with A3.6B is a very nice size.
the sovereignty topic needs more attention in general so great to see. self-hosting the model is one piece of sovereignty, but how do we handle the rest of the agent stack - embeddings, retrieval, memory, etc. Has anyone put together a practical agent stack that's 100% sovereign, where they control it all?
Anyone thinking this won't improve or is behind etc is blatantly wrong. It'll catchup within a year. Like that unknown wise and visionary man inside Google once said about their competitors: We have no moat neither does anyone else."
Congrats to the team.
For those sick of “Pareto frontier” talk, just shorthand it as “it’s the best at some very particular thing”. Obviously that one thing/tradeoff it’s good at may not necessarily be compelling, but it is either a loose sign of quality, or a sign that they’ve chased some tiny edge into the ground.
I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.
I wish nothing but luck for an EU model, but:
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
> The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content.
"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.
Was expecting that a "sovereign" AI model would at least use their own sovereign language (German) on the website as one of the options. Anyways, all the best and happy reunification day.
3.5b active params sounds cheap until you remember all 78b still has to fit in memory. curious what the smallest practical self-hosted setup looks like for german docs.
i wonder if the custom tokenizer is better in practice, the examples look interesting though
The benchmarks are impressive given the problem space they are working in.
German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.
Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?
is this like an unfortunate name clash with KolibriOS (kind of like how Google Gemini was a clash with the Gemini protocol project)?
I wonder if we can start having LLM distros: community led distributed training runs with periodic releases, open weights, FOSS code, the whole shebang. Maybe the public training sets can reach a level where an LLM trained on them can be good enough for most things, such as web search and aggregation, coding, etc.
I wonder how far we are from this. How far are we from LLM's Debian moment?
> A bigger dense model beats it. Qwen3.8 27B ...
How is a 27B dense model bigger than a 78B MoE?
I got distracted by that scroll-wheel UI component on the page. Neat!
Nice to see public goods in this space.
calling qwen 27B a bigger model.. I don't know man. My vram says otherwise.
The name really evokes strong Kotlin library naming vibes.
Interesting they recommended high end software without considering quant 4 or 8 and still used A3B which should give good throughput on cheap hardware.
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.
It still steals my IP without attribution. Now we have state sanctioned sovereign theft instead of foreign theft.
The Sega 32X game??
Aleph Alpha is just a sad joke by now. The talent isn't there anymore, they never managed to catch up to the other labs, failed to deliver on several projects, and by now are just a cash grab for the investors.
> It knows less from memory, Multi-turn tool calling is weaker, It’s not the best coding agent
Then what does it good at? Sending faxes?
does it have GDPR compliance?
The ignorant, hostile, negative, and, frankly, kinda racist comments here really are just a sad showing for the currently online crowd.
But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.
We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).
Is Germany sovereign? It outsourced its defence onto the USA. Recently Trump wanted more diesel; Germany insta-submitted, also because oddly enough Macron submitted before Germany (Macron is suspicious). Before that, Leyen committed to insta-submission with a deal that made europeans poorer (and perhaps Leyen benefits from that). Canada shows the way. Many of the smaller countries in the EU too, such as Netherlands, Denmark, Finland, to some extent Sweden as well. Every time I read "sovereign" here I have to object. Nothing is sovereign here. The whole hardware is definitely not sovereign. Perhaps some of the software is, but that's about it. Plus, who gets all the data? The big US mega-corporations sniff non-stop. Remember how Facebook sniffed Libgen and Anna's Archive dry etc..., then suddenly libgen went down. The US corporations act as huge global leeches on every step of the stair. And lobbyists benefit from this too.
> 4. It thinks in German
This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.
[flagged]
The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.