No evidence other than the fact that this has been happening steadily in all areas for many years?
You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.