Entirely agree, also things can be logically correct and well tested but not the behaviour the user intended.
Even if that's down to a bad initial prompt, or lack of data for the agent to notice the edge case I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.
You are correct that it can all be correct in one sense, and wrong in the product (“what the user wants”) sense. Of course you need to monitor the output, that’s a given. You need to question the entire system. That is the frontier for engineers.
“How is the agent lying to me?” Etc.
> I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.
I also totally agree with your response. But my point was that you can't engineer a better agentic solution. I'm arguing to never use agentic development because it is fundamentally flawed and inferior.