logoalt Hacker News

bmorgtoday at 7:30 PM0 repliesview on HN

In-prompt "security" is not reliable. You can not tell if the LLM/agent actually followed your instructions or whether it fell for a prompt injection.