logoalt Hacker News

safogyesterday at 10:00 PM1 replyview on HN

I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.

Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.

I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.


Replies

BenzeneDreamyesterday at 10:08 PM

When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes

show 1 reply