That’s a direct result of how they’re fine tuned. RLHF and other mechanisms reward quick, locally correct answers that solve the user’s immediate problem. A reward for a solution like "take a two-day pause, rip out half the modules, and rewrite the core" just straight-up doesn't exist in training datasets