logoalt Hacker News

antman • today at 7:09 AM • 10 replies • view on HN

This argument implicitly makes a few assumptions which will probably not hold in the very near future.

One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.

The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse


Replies

catlifeonmars • today at 11:43 AM

> but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds

You’re making the following assumptions:

1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)

2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

➕ show 1 reply
Vetch • today at 8:40 AM

Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too.

This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.

Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.

Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.

Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.

The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.

➕ show 1 reply
dofm • today at 7:50 AM

> One is that AI will continue hallucinating in a manner that is not easy to verify

It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.

Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.

➕ show 2 replies
Terr_ • today at 7:32 AM

> assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify

Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.

At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.

I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.

➕ show 2 replies
Retr0id • today at 7:39 AM

Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.

➕ show 1 reply
thereitgoes456 • today at 7:12 AM

AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?

➕ show 2 replies
nullsanity • today at 7:29 AM

[dead]

mathisfun123 • today at 7:15 AM

> One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.

There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

➕ show 4 replies
wolvesechoes • today at 8:10 AM

Expression of tech-faith is not intellectually honest argument.

Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.

➕ show 1 reply
pegasus • today at 8:46 AM

Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers.

Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...