Plenty of serious production projects are doing exactly that. Are you using GPT-6 Astra, or something older?
You and the person you are replying to are talking about different things.
Agents cannot be given a high level goal and then left unsupervised, for hours, without making some dumb decisions.
Astra does exactly the same sort of things that Sol or any of the previous agents do. They duplicate code, overengineer, miss the point, etc.
I was very optimistic about it when it was announced and saw all the demos, but a week later I find it only marginally better (and in some cases worse) than before.
Which serious production projects have AI agents coding on their own? And I’m assuming that means they are routinely taking tasks and deploying them to production autonomously