logoalt Hacker News

Mistral Large 4

744 points • by Philpax • today at 1:15 PM • 421 comments • view on HN

Comments

eigenspace • today at 1:33 PM

Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.

I certainly wouldnt have predicted that 10 years ago.

Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.

I'm excited to try this out today.

➕ show 9 replies
prodigycorp • today at 1:28 PM

Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.

Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.

Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.

I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.

➕ show 1 reply
simonw • today at 2:09 PM

Surprisingly it only supports reasoning "none" or reasoning "high".

That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.

The high bicycle frame is better then the none one though.

Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )

➕ show 1 reply
manlymuppet • today at 3:04 PM

Man, a lot of this discussion sounds like people cheering for the last kid crossing the finish line.

Surely we want competition and Europe involved in that, but at this point I have grown used to either frontier labs smashing the frontier remarkably fast, or open labs getting way, way closer than you would expect them to.

Mistral’s progress, regrettably, feels much slower. This model doesn’t knock anybody’s socks off. It’s (and I hate to be this harsh) mediocre, and this mediocrity has also arrived months late.

This is a pretty grim prognosis for European AI.

jakozaur • today at 1:36 PM

A strong competitor in cybersecurity as an alternative to GLM-5.3 (Mistral reports 82% on CyberGym-E2E). Visual grounding is also impressive (42% on Dense 200 vs. 41% for GPT-6 Astra).

Otherwise, behind on the broader Pareto frontier, but not by much (Vals Index: 48.05% vs. GLM-5.3’s 53.51%; $13.78 vs. $7.25 per test). Many companies will prefer it over Chinease models.

➕ show 1 reply
simjnd • today at 1:56 PM

It's a bit below DeepSeek 4.1 Flash at about twice the size. For a model that was supposed to come out a few months ago this is pretty good. Mistral catching up to the chinese open-weight models is great news. Excited to see how they will build on that!

zkmon • today at 1:38 PM

Europe needs a lot of these. Quickly. Way to go, Mistral! Keep 'em coming.

➕ show 4 replies
jascha_eng • today at 2:19 PM

-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience

That's not particularly great.

That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.

If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.

➕ show 1 reply
karannb • today at 3:40 PM

From Guillaume

> The RL run behind this preview is still in flight and shows no sign of saturation -- we will release a final version before the end of the month along with the weights of the model.

https://x.com/GuillaumeLample/status/2107461898127954001

Nux • today at 3:54 PM

Number one in Sovereign AI. Join our Discord.

lifeisloving • today at 2:02 PM

Seems hard to get customers at that price range when you're competing with open source models that are 1/2 - 1/3 the price but with similar capabilities.

People usually buy the cheapest, like Deepseek or GLM or they spend on Anthropic/OpenAI subs.. Are these models in the middle getting any users?

On a side note, I wonder if this was the popular free Space Bunny model that left openrouter yesterday.

➕ show 1 reply
apexalpha • today at 1:34 PM

Excited to hear this!

I barely use anything outside of cheap Chinese models on OpenRouter anymore. They are simply (more than) good enough for most of the things I do.

This model looks reasonably cheap. Though not deepseek levels.

Going to test it with Hermes, wondering where it will land in term of capability.

Bon chance, Mistral!

pizlonator • today at 2:35 PM

Refreshing to see this.

The pricing ($.68 in/$.07 cached/$2.09 out) makes it much cheaper than Kimi K3, GLM 5.3, and Meta Muse Spark 1.3. That's great!

But also much more expensive than GLM 5.3-flash and Spark 1.3 Contributor (the Meta-takes-your-data pricing of Spark 1.3).

So, I think it would have to be significantly better than GLM 5.3-flash to be worth it. GLM 5.3-flash is already very good.

➕ show 3 replies
Narciss • today at 1:22 PM

Le Chaton Fat is here!

➕ show 1 reply
nsbk • today at 1:31 PM

Nice! Once they make it available through their API I will be happy to support them. My local Qwen3.8 27B is serving me well, but I miss the speed and concurrency that comes with subscriptions, and I am not currently paying for any.

Tais-toi et prends mon argent!

➕ show 1 reply
mchusma • today at 2:55 PM

Something looks off in artificial analysis. Benchmarks aren’t everything, but not even close to the Pareto https://artificialanalysis.ai/models/mistral-large-4?cost=in...

I guess lots of token usage.

fancyfredbot • today at 2:26 PM

Well, with this and Beam people are going to have to stop saying that western open models are dead. This is great news. Anthropic and OpenAI may have a bit more knowledge, talent and compute but they don't have a monopoly.

Investors looking for them to make monopoly profits are going to be disappointed The premium they can extract from consumers for their models will be capped. Tokens are likely to remain close to the cost of compute, a cost which is high but falling fast.

Also, yay Europe! Although the comparison between this and mimo 2.6 is not flattering...

ianpurton • today at 2:10 PM

Off Topic - The Mistral website - Really nice design. My guess, built by a human.

➕ show 1 reply
Jeeetendra • today at 2:20 PM

49b active parameters sounds manageable until you look at the 1t total weights. what hardware does a usable self-hosted setup actually need, especially once you add a long context?

➕ show 1 reply
andhuman • today at 2:56 PM

At the end of the blog post we get this nugget. > The pace of progress from here will be fast. Stay tuned.

skc • today at 1:47 PM

We're probably fast approaching the scenario where the cheapest models will win out.

segmondy • today at 2:28 PM

I tried Mistral's last 3 large models, devstral and mistrallarge3 and the numbers were not even benchmaxxed, but just false. the models were so weak and garbage. Let's hope they are telling the truth this time around, we need more alternatives. There mistral-small and original MistralLarge models were awesome, hoping they are back!

cbg0 • today at 1:22 PM

Claims to be on par with GLM 5.3 in DeepSWE (from https://thenextweb.com/news/mistral-releases-large-4-a-1-tri...)

Luker88 • today at 1:45 PM

Mistral Large 4: 1050B, 49 Active

GLM-5.3: 753B, 40 Active

I was hoping for something that hinted at smaller models too, but I guess not.

Any competition is still good, especially now that the USA AI labs are starting to do regulatory capture.

➕ show 3 replies
aeneas_ory • today at 1:25 PM

Benchmarks are better than expected! And probably got there without distillation ;)

➕ show 1 reply
dverlaeckt80 • today at 2:16 PM

Good to see Europe is at least a little bit still in the game.

https://docs.mistral.ai/inference/model-selection-guide?mode...

Cost is stated at half the price of GLM-5.3, which is quite interesting.

Kim_Bruning • today at 2:14 PM

I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.

Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.

Not suitable for my purposes I don't think.

dom96 • today at 1:48 PM

Excited to test this on my benchmark[1], but I'm guessing it won't fare better than MiMo v2.6 Pro which is currently the best open-weight model as far as I'm concerned.

Why isn't this on openrouter yet? Is there a better router that gets models much quicker?

1 - https://bench.killswitch-lang.org/

➕ show 1 reply
valzam • today at 2:21 PM

What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.

➕ show 1 reply
walrus01 • today at 1:43 PM

The terminalbench 4.0 score is encouraging as a sign of it not doing anything "stupid" when put in a proper harness.

jrflo • today at 2:19 PM

Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.

gabe-santana • today at 3:36 PM

Amazing! just tested

rarisma • today at 1:49 PM

le chaton fat is real, my life is complete. Benches look crazy good for 1T.

mcbuilder • today at 1:25 PM

Looks like they are doing 50% off to stay price competitive with DS Flash V4.1

EDM115 • today at 3:03 PM

We actually got Le Chaton Fat before GTA 6

drbscl • today at 2:38 PM

So about 2 or 3 generations behind, just like they were a year ago?

Roark66 • today at 1:32 PM

This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.

I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.

I'd live to have one like that but EU made.

sixhobbits • today at 2:04 PM

if it's not available yet why have a 'try it today' header at all?

> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.

➕ show 1 reply
ThouYS • today at 1:49 PM

Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats

igleria • today at 2:32 PM

I thought lechonk motto was just a meme!

goatley • today at 2:36 PM

Finally, a model small enough to self-host on my 2012 MacBook Air if I don't mind my house reaching room temperature in 2 seconds.

➕ show 1 reply
tosh • today at 1:40 PM

sorting the charts like that gives off weird vibes

https://mistral.ai/news/mistral-large-4/

➕ show 1 reply
tdubey • today at 1:30 PM

Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?

➕ show 2 replies
4rtem • today at 1:32 PM

Previous one is barely in top 50 on arena.ai

staticman2 • today at 1:24 PM

Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.

➕ show 3 replies
ofirg • today at 1:39 PM

where does sit on the pareto distribution compered to Le Chaton Fat?

alpineman • today at 1:59 PM

>> Unofficially ML4, very officially: le Chonk

Honestly just nice to see a leader in this space not take themselves so seriously.

🔗 View 13 more comments