The concern is that there could have been expert knowledge input, whose importance/worth we are unable to evaluate.
I don’t believe a mathematician produced the counter example secretly, but how much did they contribute to the result?
AI isn’t magic, so to evaluate the value delta, you need to know the value of the input.
While I agree that we need the inputs to properly evaluate what this means for LLM capabilities, I don't really believe that the amount of knowledge input matters much for the overall significance of the result.
These kinds of results are interesting for LLMs because mathematicians have been working on them for decades. If the result doesn't already exist, there's no way it's in the training data, and if mathematicians have been unsuccessfully tackling the problem for decades, it is believable that the use of a new tool made the result possible, even if guided by a great mathematician.