(Background: trained, published, but still amateur mathematician.)
This is a cool blog post and I think you're going the right way, and beginning to get an understanding of the proof as you go.
I'd recommend continuing on the simplification and understanding route, until you yourself can follow the proof. Some suggestions, as I did something similar:
1. See if (or ask the AIs) if individual parts of the proof can be found elsewhere, i.e., is an argument just a copy of something else? If so, it's important to attribute this, but also this usually allows simplification ("by Theorem X", etc.)
2. Look for redundant patterns and try to combine them.
3. Ask the AI to be a critical reviewer from some journal, and try to fix its criticisms.
4. Continue simplifying! Assume that the final result may actually be relatively short.
Good luck!
> In either case I believe people who can put AI to the most value are the mathematicians themselves
The net output of math will increase, and mathematicians have more work now to unravel all this, and make it useful. AI plays the role of a monkey in the infinite monkey theorem [1]. We now need an LLM corollary - Something like: A finite number of LLM agents will almost surely find all theorems given an infinite token budget.
A wonderfully made introduction to the surreal numbers and their surrounding game theoretic concepts is this video on Hackenbush[0], a winner in 3Blue1Brown's Summer of Math competition.
[0]https://www.google.com/search?q=video+introduction+to+surrea...
> Take all the numbers you have so far. Then, “spawn” a new number in every gap between the numbers you already have (crucially, “to the left of all” and “to the right of all” also count as “gaps”). Apply this step forevermore, and you’ll get surreal numbers.
I'm not a mathematician. Can someone explain to me how this approach gets you beyond the rational numbers?
Also, this was formatted as a blockquote, but as far as I can see, this blog post is the only instance of this formulation online.
>I’ve emailed some of the mathematicians with a few proposed typo fixes, and I got confirmation that at least a few of those fixes seemed real. However, some of the problems that weren’t backed by Lean also turned out to be misunderstandings.
I think this project is really neat, but is it appropriate to cold email specialists before you've put in enough hours of effort to describe yourself as more than an "amateur"? OP's emails may have been helpful, but billions of people use these LLMs to wade into new areas and email is already low signal-to-noise.
> On the second day, there are two gaps: “between nothing and zero” and “between zero and nothing”. Two numbers spawn in those two gaps. Call them –1 and 1.
Got lost here. I think I'm officially too dumb for math.
I find the way the author communicates with the LLM fascinating. For example:
> However, I didn’t just want any result; I wanted something that pulls me.
> Initially, I asked Claude:
> Me: which unsolved problems in the Surreal Numbers research program pull you the most and why?
Note the switch from "pulls me" to "pull[s] you". What is the author's perception of the relationship/boundary between them and the LLM here?
1. Are they using it to find things it flags as interesting in hopes they might also find it interesting?
2. Do they consider "interesting" to be a universal (observer-independent) trait and are using the LLM to find things that are interesting?
3. Have they delegated their desire to find something interesting to the LLM so that it can instead find something that it flags as interesting, regardless of how the author feels?
4. Do they see it as a part of their thought process, and so do not distinguish "you" from "me"?
5. Do they see it as part of them, and are referring to the combined entity in the second person?
I would love clarification on this.
Really nice "proof guide": https://gaearon.github.io/conway-refinement/#/highlights
and "proof map": https://gaearon.github.io/conway-refinement/#/map/conway-ref...
My main thought is that he was performing something general here that is actually valuable and hard for a large proportion of humanity. It's like when Google search was a difficult thing, or troubleshooting a PC. what he is able to do here is actually a rare skill, called intelligence and he may think it's nothing, but it's actually very rare and hard for most people. The kind of judgment and interpretation of output without deep expertise is actually a very rare ability.
> a sort of epistemic performance art project.
Agreed, and it's a wonderful piece of art. I look forward to seeing the actual publication and reaction from the math community.
The Claude output in the first one-shot counterexample attempt is hilarious. I hate its writing most of the time but this stuff is next level deep-fried slop.
> And the control column confirms the resonance-necessity conjecture empirically: break the skeleton alignment and the joint kernel dies at the constrained window, exactly as the transversality heuristic predicted.
> The den has air in it.
> Drift fuel exists.
My experience has been similar; I find ChatGPT to be much stronger and more precise at math and in communication. I also can not do better with a multiple agent flow than I can with a single agent.
This year, LLMs have been involved in a number of interesting proofs of conjectures. But that is not even half of mathematics. Has anyone tried to use an LLM to generate a mathematically interesting conjecture, on the level of Conway's refinement conjecture? If so, what happened?
With all the talk of mathematicians possibly being obsolete, I'm wondering where the future conjectures that future LLMs would prove might come from.
Damn, I was not expecting Dan Abramov when I read the title.
I find this whole post fascinating in the context of https://news.ycombinator.com/item?id=49738091 and particularly this excerpt from Gowers:
> Instead, I have a more complicated view, which I actually expressed in my essay The Two Cultures of Mathematics a quarter of a century ago, and which can be summarized by saying that there is a spectrum of attitudes in mathematics to the relationship between problem-solving and conceptual understanding. At one end of the spectrum you have mathematicians who are primarily motivated by the wish to solve problems, who see conceptual understanding as a very important means to that end. At the other you have mathematicians who are primarily motivated by the wish to attain conceptual understanding, who see problem-solving as a very important means to that end.
Before, understanding and problem-solving-ability were so interdependent that distinguishing between the two was practically very difficult and probably wouldn’t have changed anyone’s research agenda. Now, they’re not connected, and this guy just did the ultimate meta-experiment of seriously undertaking a project that is intentionally 100% problem-solving and 0% understanding to prove it (maybe 99% and 1% but pretty close. In his transcripts, he never asks ChatGPT about the math, only about its opinions of the math).
As we (as a society) sit around asking ourselves what mathematicians (and software engineers, and anyone in deep technical fields) should be doing all day, we now have this case study to show us how wide our range of options has become.
Someone please vibe-prove that ZFC is inconsistent.
That's awesome! Congratulations!
I'd imagine that in three months when we all have access to communicating agent swarms this should be easier
Perhaps one thing you should devote effort to is ensuring this has not already been proved in the literature.
this is a great blog post, documenting a very real process of what it's like to create large results with fallible models
though I am an expert at coding, the author's process sounds very similar. constantly double checking, asking for explanations, having AI adversarially check its own work, trying to detect bullshit
free time spent talking to llm, what an achievement!
Here's my conjecture. Large Language Models are the great filter. They represent a local maximum in the technological advancement of a species from which we will not escape.
I don't know why the author could claim this is "their" proof, and they kept saying "they" did this, "they" built that. but in reality everything is done by the LLM and the author is merely asking it to do things. i guess they did contribute money at least...
> Me: btw how’s your mood overall?
LOL. mood??
[dead]
[dead]
“Guy amuses himself in token burn rpg while simultaneously demotivating anyone who would have found value in this process.”
For some reason, this approach makes me think of the difference between “wizardry” and “sorcery” in some fantasy magic systems. The magic of “wizards” is fundamentally based on a deep study and understanding of arcane things, perhaps assisted by some (necessary or helpful) tools of great power. “Sorcerers” summon supernatural beings and are able to control them, cajole them, and protect themselves and others against them (with more or less success)... but the actual desired magical effect is performed by those beings.
Computing has historically been a field of wizardry. It's... interesting (?) to see so many people pushing so hard in the direction of sorcery, and in fact applying that sorcery to other fields, in which they themselves aren't quite able to validate whether the spell worked or not.