I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline.
I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they sell the same public models to their private customers (military etc). We also know they talk about "unpublished internal models" for things like the last HuggingFace hacking incident. So it's not a bad theory.
This is kind of my intuition as well.
I suspect that the models we don’t see are decidedly better than the models we do see.
> I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible
How could we really know how much "compute China has" in reality? Is it possible that whatever estimates people has come up with for both China and the US might not be 100% accurate?