Once it's open weight people will be able to inspect and compare it's tokenizer, architecture etc and tell.
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.
it's extremely unlikely that they re-use anything from a chinese model, that would be obvious quickly, what's more likely is using documents produced by a better model to create synthetic data.