Its quite interesting to see that at least the early days of AI so far have not been a winner-take-all runaway acceleration game where catchup is impossible.
I certainly wouldnt have predicted that 10 years ago.
Very glad to see Mistral still in the game even after some big stumbles with Large 3. I deeply hope that this model is 'good enough' that it becomes the European go-to, giving them the resources to keep the pace up.
I'm excited to try this out today.
> have not been a winner-take-all runaway acceleration game where catchup is impossible
From the Mistral site:
> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
It is pretty capital intensive!
For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.
But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.
We haven't reached RSI yet. Once any entity reaches RSI, the runway scenario will happen.
I strongly disagree with this "early days" framing.
AI is an idea 60 years old. We are on the 3rd or 4th generation of AI development. Three years into the current iteration of products.
This is not early days by any measure. LLMs are a result of a very, very mature research field.
What do you mean "good enough"? Did you mean "large enough"? ;)
Disclaimer: I'm not sure how much of an IYKYK factor applies to this joke.
Yes, so far the competitive dynamics feel more like cloud computing than web search.
Seems silly not to have predicted that 10 years ago. I feel like it's long been obvious that smarter models being available will mean way easier cheap synthetic data and access to tools that will speed up competitors as well as consumers.
I think it's a mistaken belief that AI as we found it is the exponential runaway train.
So it makes sense, since all you need is compute, that there's a ceiling and specialization is going to be more valuable then some super AGI.
Especially since the worst people seem to be the ones who think they'll all run away with the bag.
How do you figure? I haven't met a single person who doesn't use Claude or Codex for programming in any serious way.
I mean, Mistral is about 9-12 months behind here when you look at its overall benchmarks versus the models released around a year ago.
On a purely technical level, maybe? But in terms of actual revenue, is there really any chance of anyone catching the big labs?
Obviously, this is only a valid question if you don't believe that open weights are about to eat their lunch and their revenue is about to collapse, or they're running a super unprofitable ponzi scheme propped up by investor money that's about to collapse like a house of cards. I don't find those positions credible at all though.
If you do, then this question isn't really for you, as I'm more interested in thoughts from those who think that OpenAI and Anthropic in particular are about to be the largest companies on earth in a couple years. Could anyone catch them at that point?
[dead]
I think a big part of that is the Chinese publishing the solution for everywhere hurdle in the road they've encountered in the form of a paper.
Deepseek essentially releases instruction manuals in paper form.