logoalt Hacker News

rcxdudetoday at 11:04 AM1 replyview on HN

I would not really call this a prompt injection attack, since it doesn't really hijack the agent to become malicious (something the article does discuss later on). It's more a trojan that's aimed at tricking Claude specifically.


Replies

bjackmantoday at 3:00 PM

Yeah I jumped on this quite excitedly but it's not prompt injection at all.

To be fair to the authors they don't actually say it is. But then they contrast it with the "0.00% prompt injection attack success rate".

The upshot is kinda the same - this is still evidence that we should be sandboxing our agents. But it doesn't actually challenge Anthropic's "our models are too clever to prompt-inject" vibe.

show 1 reply