just curious, how do we know that le chonk isn't just a fine-tuned chinese model? and/or distilled from US models?
It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe".
https://mistral.ai/news/mistral-large-4/
Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell.
It's in the announcement: "ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own datacenters in Europe".
https://mistral.ai/news/mistral-large-4/