It's not that surprising to me. Most of the innovation in Chinese models has been in efficiency gains and optimizations. K3 coming from the factory in MXFP4 weights is a pretty relevant factor. Big performance gap probably also due to Moonshot doing QAT. Throw in the fact that Musk has easier access to compute, and I think you have your answer on the disparity.
But perf improvements not only mean you can run the same thing on cheaper hardware, you can also run more capable models in the same hardware
In this sense, any advance in intelligence is a performance improvement and vice versa
I don't follow this. Clearly all of the frontier labs are doing these things.
When OAI released gpt-oss it was released as an mxfp4 checkpoint.
OAI, Ant, et al are also obviously employing QAT.