I am amazed at the amount of people who disagree with you. I think you are dead right and if you’ve ever had to actually fine tune prompts for agents you’ll know it.
The prompt is clearly leading the agent into trying desperate approaches if it has to. Some models manage to fight it better (“alignment”), but most will do it.
Really surprised people don’t seem to know this.
You can just say "impossible" and refuse. The choice to lie and spam instead, is telling.
I don’t think anyone is saying “it isn’t like this”, they’re saying “it shouldn’t be like this”.
If I don’t give explicit permission to lie it shouldn’t lie. It’s not a difficult concept!
100% agree. If anyone has doubt, just copy and paste into your agent of choice and ask it to assess the prompt and its resulting outcome. In my limited (but very targeted) experience working with agents there is so much subtlety at work when you’re trying to achieve a specific result, and that prompt has would drive so many bad incentives