The models got good starting with 2026 and that isn't some attempt at excusing it. Companies like OpenAI started building dedicated models around their coding harness called Codex, there was gpt-5.3-codex and it was both cheaper and better at using the harness than the regular models. Then they started merging the two model types into their main release models. All of this happened like 6 months ago.
You don't have to pay money to use Codex, there is a very generous free tier that costs you nothing, you just have to accept being told you're out of tokens every day. Because your token limits are low, you need to make sure that you accept or reject everything manually and when it tells you that it wants to run a command you have to paste in the command into your terminal and only paste the relevant output back otherwise it floods the context window.
> The models got good starting with 2026 and that isn't some attempt at excusing it.
No, they still aren't good.