> A challenge with this kind of study is that coding agents (Claude Code, OpenAI Codex) only started working really well in
2026? 5? 4? 3?
Heard this one way too many times.
Nobody was saying coding agents started working in 2023 or 2024, because the category was defined by Claude Code which was first released in February 2025.
Yep, the goalposts just keep shifting. In reality: they still don't work well, unless you're content with producing low quality work.
It seems to be true this time though; I have observed it myself and heard it from several experienced developers I personally know and respect. It feels like some threshold was crossed with Opus 4.5 and Gpt 5.3, where the models are now able to reliably solve certain classes of problems that were previously unreliable.
Time will tell of course, and it’s early, but inflection points do exist with progress.