It's just model size and heavy RL, sometimes they overfit on specific tasks. RL can get you very far, prior models did not have such a focus on RL for agentic setups.
Look at deepseek, they improved it just by doing a lot of RL and you can see it from how it behaves. You provide very little information about a task, but since they are trained on similar tasks, they come up with a lot of assumptions and details on their own, because they were trained with such an info during RL.