Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.
We weep for Dario, that he had to suffer such a devastating attack against his Terms of Service.
I have very little sympathy for thieves who get robbed of the goods they have stolen.
If you understand what they have achieved here, then the notion that they are bottle-necked on training data is absurd.
I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?
Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?
Eh, even if this was true, then they're merely stealing from thieves. Anthropic did break a ToS or two to get training data themselves.
[dead]
Anthropic infringed the copyright of basically every author on the planet: https://www.anthropiccopyrightsettlement.com/
No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.