This is tricky, because we really want language-independent training of skills. We know that self-play type of reinforcement learning is incredibly effective when possible. But at the same time, they are our tools - so we need supervised language training for this reason? It's possible that training just needs to be rebalanced so that RL with rewards is balanced with rounds of language adjustment. And to really make that happen, benchmarks need to score the models on that.