logoalt Hacker News

himata4113today at 8:05 PM2 repliesview on HN

It's not that their strategy is to train smaller models, it's the only choice they have. Training SOTA takes anywhere from 1.5b to 150b. We don't know the real cost of training for the chinese models, but mistral neither has the compute nor money to do that.


Replies

lucrbvitoday at 9:08 PM

Mistral has the capability of training such models. Take a look at Poolside[1], they are claiming to pre-train their Laguna series of models on 4,096 NVIDIA H200 GPUs[2]. Mistral has approximately 13,800 NVIDIA GB300 GPUs, which are nearly 2x more efficient for training.

The problem with Mistral is that they do not seem to have aligned incentives to train big open-weight models, even if the teams would like to.

[1]: https://poolside.ai/ [2]: https://poolside.ai/blog/introducing-laguna-s-2-1

maelitotoday at 9:00 PM

Do you have a reference explaining these costs ? Part by part.