I suspect that because each RLVR episode injects ~1 bit into the models capabilities, and training on a reasoning trace injects ~megabyte into a models capabilities, distillation is powerful enough right now that they’re all basically the same model