logoalt Hacker News

rmunn • today at 1:17 AM • 1 reply • view on HN

Short version of the article: no, not even close.

Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.

My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.


Replies

janalsncm • today at 1:27 AM

A bit too pessimistic imo. I agree that AI can’t automate things end to end, but a good deal of R&D involves kicking off a training run and babysitting it.

If your training run dies at 1 am and you’re sleeping, you won’t find out about it until the next day. You can lose up to 18 hours of work depending on when it happens. Based on the error it might be as simple as tweaking a single hyperparameter and rebooting, which is something LLMs are usually capable of.

Even just that task means I can kick off multiple runs over the weekend and have confidence they’ll finish. It’s a game changer.

➕ show 1 reply