Tests are inadequate to catch bugs. Tests are like putting thin net in some little sections of a window, but property based testing is like something more like a decent net but not a solid barrier. And agents will lie, misunderstand or hallucinate things when asked to "prove it." And none of this fights off the bloat, performance issues, and decrease in maintainability. People keep thinking the process will save them from actually having to think through and put things together correctly. It's too bad that people are missing the satisfaction from personally building solid, correct things. And users are getting increasingly worse software because of this trend.
Entirely agree, also things can be logically correct and well tested but not the behaviour the user intended.
Even if that's down to a bad initial prompt, or lack of data for the agent to notice the edge case I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.