logoalt Hacker News

HarHarVeryFunnyyesterday at 10:21 PM1 replyview on HN

> What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one

If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.

Does this aspect really matter? Not really, other than OpenAI wanting to present this as all the work of their model.

**

https://cims.nyu.edu/~tristanb/statement.pdf

I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.

I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.


Replies

famouswafflesyesterday at 10:38 PM

>If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.

As it seems and as they tell it, they started the run modestly and diverted more resources towards it as it looked more and more promising. The run didn't start with 10k agents for instance. The point is there isn't anything humans are doing in this timeframe against all this text that would count more than "little human output". It's still a fair assessment I would say.