logoalt Hacker News

aabhaytoday at 8:09 AM6 repliesview on HN

My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I want to know:

1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?


Replies

whattheheckhecktoday at 2:12 PM

Yeah I remember reading about something along the lines of Mathematics is now about the scaffolding around you find the problems/solutions not just the problems and solutions. For teaching purposes. This was before this ai craze

einpoklumtoday at 8:19 AM

Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?

Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.

show 4 replies
dist-epochtoday at 10:43 AM

I don't think you want to bring cost into this argument.

Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.

Do you really think that if you paid that to humans, they will deliver the same results?

show 5 replies
irthomasthomastoday at 11:26 AM

[flagged]

simianwordstoday at 8:34 AM

There are people who can’t grasp the universe without mandatory randomised controlled trial. Would tomorrow be a Sunday? Need an RCT for that boys!

My point here is to not snark. But there should be some level of self skepticism that doesn’t warrant an RCT theatre.

show 2 replies
azan_today at 12:04 PM

> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.

I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.