Can someone who understands maths more than me explain why it could only solve 372/8000 problems?
What was it about the other problems that made them unsolvable? Was it just a time constraint, or are they just harder problems?
If it took 3h for one of them, perhaps there was a time/compute budget cutoff along with a sorting based on some relevance.
My guess would be that these particular problems were vulnerable to an attack which built on recent advances and potentially tied in something unexpected from a distant area of mathematics. "Harder" is becoming harder to define. Harder for humans is probably not harder for LLMs.
There must be an element of luck, if they ran the remaining problems again with the same time constraints presumably a bunch would be solved
I'm a research mathematician. From what I can tell, the answer is roughly comparable to: if you posed 8,000 challenging open problems to the human math community, you might expect to see 372 of them solved within five years.
Probably some combination of: some of the 372 problems were easier than the rest; the AI got lucky on these 372; there were existing papers out there in the literature which proved especially helpful for these 372; and other similar factors.