logoalt Hacker News

antirezyesterday at 9:35 AM6 repliesview on HN

So many mathematicians over the years tried hard and failed, but now Anthropic just for some PR magically did it? And this after LLMs obtaining different math wins? What is your logic here really escapes my understanding.


Replies

glimsheyesterday at 10:03 AM

The parent's absolutely nonsensical post highlights how polarized AI (as everything else) is today. I can understand someone being opposed to AI on moral, cost-benefit or productivity grounds. But we're seeing a lot of extremist "AI is good for nothing" posts out there nowadays.

show 2 replies
raxxorraxoryesterday at 11:21 AM

Those silly advertisers do everything for exposure and if that means digging yourself into a niche alleged mathematical theorem to refute it, it is what needs to be done!

Of course it would be really interesting how Claude approached this. Probably with some constraints regarding the input. And it would be interesting what these constraints were.

show 1 reply
YeGoblynQueenneyesterday at 10:38 PM

I mean we have no idea what happened exactly, how Fable was used, how many times it was run, whether earlier models were also tried, what was the prompt, how long it run for, etc etc. All we have to go by is a tweet.

Why not be skeptical about that?

show 1 reply
jibalyesterday at 10:13 AM

They gave their "logic", such as it is ... and it's utterly irrational.

Note that the "they" who published the counterexample on X is some rando mathematician (Levent Alpöge) working for Anthropic, not Anthropic the organization. He posted the counterexample in a tweet -- reason enough for "not disclosing the LLM chat session". There's no reason to think that it won't provided if asked for, but it hardly seems relevant.

show 3 replies
0x1ceb00dayesterday at 10:32 AM

[flagged]

show 2 replies
William_BByesterday at 10:35 AM

There's a big difference between one shotting a counterexample using AI and using AI to find a counter-example by brute-force.

Both are impressive, of course, but they're hardly comparable.

show 1 reply