I can tell you it was not correct in January of 2026. The real swift started happening with the latest models opus 4.7, fable 5, kimi k3, glm 5.2.
That's when the models started to be coherent enough for real work.
They still fuck up, but it does not feel the code was written by drunk interns anymore.
[flagged]
A month ago I was told January of this year was the inflection point. Month before that the inflection point was December of last year. I'm not saying the tech isn't getting better but are the fundamental limitations being surpassed or are the long tail failures just being pushed further away? Because if its the latter this game of "well models really got good nine months ago" won't stop.