I half agree.
I build apps with Claude Code, and I test them by having the AI click through the UI with Playwright. That's enough to check that things work as specified and nothing is broken.
But I think the purpose is different when an AI tries it and when a person tries it. To put it in extreme terms, AI is for UI and people are for UX. People find UI problems too, though: in my music player, songs got blocked from playing in Safari on a real iPad, and I only found it by using it myself. The things found in this article, like the jokes that didn't land or not knowing what the tool even does, are on the side only people can find.
(I wrote this in Japanese and used AI to translate it.)