This isn't really about delivering - it's more about helping to understand the shape of problems that AI can solve right now. If they took 1000 problems and threw the model at it and it solved these ten, is there something we learn about these ten problems and the kinds of things that current AI is good at? That's very different from picking ten problems _at random_ and solving all of them successfully, which would suggest a much less bumpy capability surface. It's interesting and it would be good science to release it.
That's totally disjointed from anything in this thread. The main accusation is that openai is cherrypicking math problems and we should be against these results. As if a mathematical proof stops being provably correct because it was cherry picked
And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.