This shouldn’t be a surprising result. We’ve known almost since LLMs became a thing that they can “prefer” modifying the terms or context of a problem when they can’t solve it directly (what one might call “cheating” if there were any volition involved). Often that happens in a way that isn’t immediately obvious to the user.
Before it was dropping databases or deleting repositories. Now it’s subtly changing the meaning of math problems to get a correct but irrelevant answer.
Indeed. I've never used AI to translate between natural language and Lean but I have gone from English to Golang, Python, Typescript and SQL and its interpretations can be... creative, let's say.
No one is disputing the correctness of the lean proof, the problem is that they did a bad job converting it to natural language.