First few models will always be slow improving and worse. The way to improvement is working your way through a gajillion evals [1], finding bugs, gaps, and curating training data (this part involves human design as well as raw inference compute) to fix it. This is very time intensive and can't easily be "done once and then everyone has lesser work to do" since every model is different. Well, one way to accelerate it is to simply have more compute, which mostly openai and anthropic have[2].
This is mistrals first 1T-scale model and I expect the 4th or 5th generation to be close to the best for many purposes.
[1] These evals differ from the public ones like terminal-bench, are sometimes model-specific, need real, diverse usage to actually create, and are held secretly since quality of eval is the first driver behind the next step improvement of a model.
[2] It is not close. This model was trained on less than 4k GPUs, whereas astra used north of 100k GPUs.
Mistral is not that new a player though. How can we give them this much grace when other players like xAI have done more in even less time? I don't think coddling Mistral helps them.
And to the point of scale and training cluster, so what? Not only do Chinese labs have smaller clusters with less empowered GPUs, compute is Mistral's responsibility. You can't take away from other labs just because they fulfill that responsibility better.