logoalt Hacker News

wg0 • today at 1:41 PM • 3 replies • view on HN

No it is not. Only maybe for the noobs or vibe coders.

People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.


Replies

balder1991 • today at 4:21 PM

I’ve been saying that. When you have no idea what you’re doing, you *need* the latest greatest model because it’s the only way to reduce errors.

For people who have some expertise, the models accelerate the grunt work, but you’re the one validating it.

CharlieDigital • today at 1:53 PM

I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.

Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.

Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."

Yeah, not that smart.

➕ show 1 reply
AIblemblio • today at 3:09 PM

Try a bigger code base or more complex stuff and you will easily see that the solution, speed and amount of problems Opus5.5 solves vs older models is relevant.

➕ show 1 reply