logoalt Hacker News

a1oyesterday at 8:54 PM1 replyview on HN

My experience is Codex is much better but less creative. I use its agent through GitHub Copilot or the agent interface from JetBrains. Try the GPT Sol 5.6 on medium, it gives me good results


Replies

weitendorfyesterday at 11:50 PM

My experience with Codex is that it goes off to do its thing for 10-60 minutes and either nails it and comes back with everything done, or comes back with something that I almost can’t believe a near-SOTA model would think I wanted based on my prompt, or is of acceptable quality.

I think the tradeoff to Claude being so needy is that if you let models just run away with an inaccurate or incomplete understanding of what to do, they can go really far off the rails AND spend a lot of time/money doing it AND come back with something that literally doesn’t make sense or doesn’t work.

I prefer dealing with Claude’s reliable cringe to the aloof model that tries to play it cool when it needs help.