When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
Are we missing the forest for the trees here? If a math problem falls in the forest but nobody is around to understand it does it make a sound?
How can we possibly make use of these breakthroughs if we don’t understand them? How could we ever make anything useful with them?
Are we ready to just let go of our intellectual faculties and give them to a giant supercomputer nobody understands? How do we tell truth from fiction?
I get the impression that the value of unproven conjectures is more in the new math and techniques that may be discovered - by humans - trying to prove/disprove them, rather than much utility in any eventual result.
Take something like Fermat's last theorem - I'd be curious to hear of any use of the result itself, but there was a massive amount of new mathematics generated by those working on it, whether ultimately successful or not.
These AI math proofs are interesting testament to the power of reinforcement learning applied to math, obviously reflecting the axiomatic self-consistent nature of math itself, but it doesn't seem they have the same value as a humans working on these problems since they are using known math to solve them rather than inventing anything new.
However, it would still be interesting to analyze the LLM lines of reasoning that lead to any of these results, since there may be value there even if no new math, just as human Go players have found value in analyzing computer Go.
Still, as Demis Hassabis has himself said, the real goal with AI is discovery and creativity - you want to create the thing that could design the game of Go in the first place, not just play it. Similarly with math, while there is interest in seeing an AI "play math" using the rules of the game, what would be of much more interest is the AI that can create new math, in the same way as Andrew Wiles did while proving Fermat's last theorem.
> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
Math is an incredibly broad field. I mean, you don't expect a traffic engineer to understand anything about nuclear reactors, do you? Yet, they are all 'career engineers'.
This seems like an extreme exaggeration from a few people claiming to not understand a very recent result. This cycle of a new result being discovered, and reaearchers needing some time to truly digest and disassemble it, is normal.
The field has been dealing with this for a long time.
Since at least 2014, which was my first brush with the phenomenon when someone published a 13GB proof [0].
The consensus is that such a proof is potentially illuminating, though further work is likely required. If for instance, conjecture A is true if and only if conjectures B & C are true, and B is proven false through one of these such proofs, then we can see that A is also false given that we accept the disproof of B.
Though, the sense is that further work is likely required because it is easy to see that further work along the same direction, or in directions depending on the proof will be hard or impossible if there are not enough humans or agents that are capable of understanding and utilizing the proof. Making it 'more elegant' will increase it's utility despite not proving anything new.
This is adjacent to all of the work done to create multiple proofs using different techniques. Having the same information (that X is so) in different languages (algebraic, geometric, via harmonic analysis, etc) allows for researchers not familiar with the original technique to participate in further research.
[0] https://www.newscientist.com/article/1997488-wikipedia-size-...
I think it's the scientist version of "I vibecoded ten apps this weekend (at one point I'll have real users too)".
Because it's math, it's all mysterious and genuinely impressive, but in the end, if no human cares about it (apart from attention grabbing "it's so over" tweets and articles), does it really matter?
Even the experts on the problem being solved find the writeups nearly impossible to read.
Example: https://nitter.poast.org/henryquantum/status/208362369543662...
Seems like a disservice to the community that openai put so little effort into producing good writeups...
Latest presentation of Terence Tao on what current advancements in AI mean for math discusses (among other things) those issues: https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.p...
I don't mean to be dismissive, are these just old puzzles with no practical use whatsoever?
Mathematics is an unusually dense (if not the most dense...by a few large steps) field. So lots of areas of mathematics are extremely deep and narrow without any real shortcuts, even for seasoned mathematicians.
Your observation is perfectly on point, I think the season of companies announcing breakthroughs might be over soon (unless they somehow manage a major achievement, P vs NP or similar). At the same time, mathematicians will be left with superintelligent machines solving the actual math for them, much like software engineers nowadays. This was unexpected, and unexpected at this scale up to a couple of months ago.
In college the punchline for all the engineer, physicist and mathematician jokes were something like "The mathematician says: Yes, there is a solution."
In all serious "I don't understand any of this it's way over my head."
My perception is different. I start from the axiom that AI is a massive compressed corpus of knowledge. That it can find solutions suggests to me that the solutions were already known, but simply lacked publicity. This is less about discovery and more about pattern-matching.
I find some interesting parallels in chess, which often has lots of analogs with math to begin with. But chess went from a pure human endeavor to one where supercomputers aided by world class players finally managed to eek out a slightly suspicious win against a world champion (approximately where we are now in math) and to now a days - where your phone could easily crush the world's strongest player, who is also probably the strongest player of all time.
The way the chess world adapted this was initially to try to understand the machine. After all chess, like math, is complete information - so you can easily see the computers 'thoughts' in terms of the exact moves its saying are best in a variation and how it might respond to any other idea. But it quickly became clear that this wasn't working so well.
Players would regularly get positions that the computer says 'and black wins' and then proceed to lose it convincingly, simply because the positions were so extremely weird and difficult to play that even if it might be technically winning, it's the sort of position where you're walking a fine line with lots of complex moves to find. Humans aren't computers and even the best of us can't play like one in weird positions.
Now a days they're taken more in balance. The computer's evaluation of a position is probably about as good as you can get, but playability matters much more in practical terms. Knowing the eval of a position doesn't really matter if you don't understand the position. Knowing the answer can help with understanding (for instance computers have radically reshaped and improved human understanding of space in chess as we noticed computers obsessing over it) but I think the days of 'oh the computer says it's winning, so I should be able to take it from here' are near to gone.
There's no honor in being stupid, yet I must admit I am stupid, as there's even less honor in being stupid and pretending otherwise.
A lot of people were super hyped about OpenAI's 10 discoveries, but I still don't understand what they mean, and even if I did, what are the implications.
Like, what are non-sofic groups, and what follows from the conclusion that they exist?
I mean in the sense that quantum mechanics might make my head spin, but it's because of that that we have stuff like semiconductors, which have been one of the most significant discoveries.
The Fourier transform is one of the reasons we have fast telecommunications and radars.
What practical things are possible or might be possible due to these results?
i doubt anyone really wants, or cares, to understand what the AI comes up with. If it's not related to your work or your name isn't associated with the discovery then I would think time is better spent on things that are.
This has been an ongoing debate since the computerized proof of the four color theorem fifty years ago.
They claimed them as advances, not breakthroughs.
> When OpenAI posted about their 10 breakthroughs, I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
Huh?
Mathematics is an extremely wide subject, it's perfectly normal even for two professional mathematicians not to understand each other's work. Have you considered that maybe they just don't work in that area?
Are you implying OpenAI's paper (which was, by the way, edited by humans and provided Lean certificates for most of the proofs) is actually gibberish? That's flat-earth levels of conspiracy.
>If a math problem falls in the forest but nobody is around to understand it does it make a sound?
This is not new nor unique. There is plenty of research, especially in math, which can really only be understood by a few people in the entire world. It is not uncommon for a proof to be presented by a mathematician which, initially, is only understood to that mathematician, and it can take a long time for even another mathematician who is an expert in the same field to be able to confidentially say they understood it.
>I saw lots of career research mathematicians say things mostly along the lines of “I don’t understand any of this it’s way over my head”.
This is meaningless in a vacuum. If you give a novel proof in some niche subfield of topology to a competent mathematics researcher who focuses in number theory, they'd say the same thing regardless of if a human or machine wrote the proof. A bunch of "career mathematicians" on Twitter proclaiming this doesn't mean anything other than these people aren't currently equipped to understand the contents of the proofs. That's fine and normal, but the idea that anybody with a PhD in Math should be able to pick up one of these proofs and give it a skim and be able to say, "ahh, yes, quite clever, it seems so obvious in retrospect," is absurd. That's not how this kind of research works.
>Are we ready to just let go of our intellectual faculties and give them to a giant supercomputer nobody understands?
Nobody is blindly accepting these proofs as valid. ChatGPT isn't spitting out a wall of text and proclaiming that they've solved a previously unsolved math problem while everyone is saying, "well if an LLM says it, it must be true!" lol
These proofs are being checked by automated systems (which have been in-use well before LLMs have existed) as well as being checked over by actual experts who are actually capable of (and motivated to) verifying these proofs. But that work still isn't done. There's enough evidence that these companies are confident in saying these proofs are correct, but there's going to be a lot of ongoing work from people to continue to verify and, more importantly, understand these proofs. It's literally some of these people's full-time jobs to do this.
>How do we tell truth from fiction?
When was the last time you verified even a classical, relatively simple mathematical assertion? How often are you just relying on a larger system of experts to ensure that we're not just blindly accepting fiction as truth?
That's not to try to stick it to you personally, but it's just highlight that there's an entire system in-place here that you're not aware of and that you don't have an understanding of that is working just fine including in this context. Real mathematicians aren't going to lazily start letting OpenAI assert whatever they want about their products solving these kinds of problems without heavy scrutiny.
Maybe the "beautiful, elegant" math is really just accidentally that way, just the tiny cross section our dumb human brains can understand. The vast majority of it could be inscrutable, ugly, chaotic and seemingly meaningless.
I have the same questions about human mathematicians!
I can't tell if xkcd #435 is still true, or if math is just as mushy as everything else seems to be. When a math proof can only be understood by a handful of people, what does that mean about that proof? I think the LLMs are pushing a problem that existed already and pushing it further.
The field of mathematics is smart about this and knows the difference between a pile of Lean code and understanding, and mathematicians try to get from the unintuitive explanations to something that makes more sense, e.g. https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the... (where incidentally Tao used a chatbot to help take apart the problem, but with a lot of interaction and work from his side).
Making things understandable is mathematics, and more generally a kind of intelligence, and is crucial to continued progress. You couldn't use algebraic geometry to disprove a conjecture if people hadn't organized (what could have been just) a pile of random observations into something called algebraic geometry.
Historically LLMs have done best where it's possible to train using an objectively verifiable reward function. Computer programs are pretty good on this front and so are Lean proofs. (Of course, they don't only do things you can RLVR heavily, but those have progressed fastest.) Not sure where 'making mathematical knowledge more understandable' falls on that spectrum.
And understandability isn't only a thing for advanced math. Keeping computer programs from becoming a mess is a challenge in high-level organization too, and the chat with the user is an explanation task. If you look online at some of the stuff people say about large LLM-built codebases (SlopCodeBench is a neat effort to make make it concrete, but common wisdom seems to mostly agree on the general problem) and what people say about chatbot prose, I don't think everyone considers those solved problems!
It's hard to tell how thoroughly the labs grasp and care about this at an organization-wide level. I'm sure at least some maybe-results exist inside labs but haven't been published because the humans couldn't verify them and didn't want to be embarrassed with a false result. (Maybe also why counterexamples are a lot of the first results published: often simple to verify, even if hard to obtain.) A good sign would be if results in a few months come out more like what mathematicians consider well-written papers explaining results in a more intuitive way, fewer shocking announcements of bare counterexamples in tweets. It's probably a slow climb to get there.