You even notice that with the recent opus and fable models by Anthropic.
If you give them a wide open problem statement, they'll start talking a lot of semi intelligible gibberish.
My guess is that this happens because that's not what they are evaluated on anymore for these kinds of tasks. The generated code is evaluated (in this case the lean code). So talking a bit of gibberish in the language part so you have more test time compute is not punished.
We need to recognize this as a failure in training. It did some useful stuff but it can be much better. A training signal is likely missing.