logoalt Hacker News

simianwords • today at 11:40 AM • 1 reply • view on HN

> When I task my agent a hard problem and it comes back with a 5000 line PR that I don't follow I reject it and work it into a better shape. I certainly don't slam it into production because it has a 'high signal' of the program I asked for.

??? That's literally what everyone does. You are creating a hypothetical that is nonsensical. Here's your hypothetical:

1. Agent gives you 5000 lines of slop

2. You reject it and just do it yourself

This is reality

1. Agent gives you 5000 lines of slop

2. you realise that it has done a lot of research and is mostly in the correct direction and you ask it nicely to refine it

3. verify that you understood it and push it to prod

Are mathematicians babies that they need a completely different approach?


Replies

samuelknight • today at 2:25 PM

The mathematician in this case didn't ask an AI for the proof, they were handed a badly written artifact and other people are messaging him asking for analysis. He has no visibility into the process that produced the paper because it came from an internal OpenAI model. The bad behavior here isn't coming from the complaining mathematician. OpenAI hasn't sufficiently developed their workflow to create cutting edge research AND publish it intelligibly, which is odd because the latter should be easier. OpenAI has perverse incentive to do this because spamming badly written papers can still give them priority credit for solving open problems at the expense of mathematicians (who are currently indispensable because they decide whether claims are credible). What Mr. Karagila is trying to enforce are the old standards that review should come after refined explanation of the solution. If that was the case before AI, I can't see a reality in which that isn't the standard now that papers can be pumped out at faster-than-human time scales.

➕ show 1 reply