According to Scott Aaronsson, OpenAI set their agents on 8000 different problems, and got 372 final proofs. Even spending twice the original effort on simplifying and reworking those proofs so that they do not "feel like something written by someone who’s on psychedelics" would only increase the compute by less than 10% (assuming all the agents had a similar token budget).
The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.
According to Scott Aaronsson, OpenAI set their agents on 8000 different problems, and got 372 final proofs. Even spending twice the original effort on simplifying and reworking those proofs so that they do not "feel like something written by someone who’s on psychedelics" would only increase the compute by less than 10% (assuming all the agents had a similar token budget).
The fact that they did not do so can only mean that either (1) their agents currently lack the capability to do it, or (2) OpenAI are completely indifferent and do not care in the slightest if the proofs are understood or not.