But they both spent tons of money on data collecting/labeling/generation, how is it bad compared to distillation? I thought their data are much better if they spent that much, and it seems they are stupid because with that much of resources putting in there with merely no output compared to the frontier models.
Creating a RL example by hand is hundreds of times more expensive than generating one using an LLM.
Of course the Chinese companies have incredibly talented researchers, and smaller, better organized org structures which account for the rest of the difference.