> have not been a winner-take-all runaway acceleration game where catchup is impossible
From the Mistral site:
> ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe.
It is pretty capital intensive!
I’m pretty impressed that they managed to get that close to the frontier with such a small cluster!
According to Grok thats 7-10 MW. Tiny numbers.
To put that into context, the last wave of capacity SpaceXAI added 400-450 MW.
That’s kinda very small and light for modern trillion-param LLMs.
These cards are like $3k each? That's, what, $12M and you keep the hardware? Honestly doesn't seem too bad.
That cluster is literally orders of magnitude smaller than the compute pools used by Anthropic or OpenAI.