logoalt Hacker News

mnicky • today at 1:06 PM • 1 reply • view on HN

> This is not prompt injection. This is a prompt entered by a human through the Claude UI.

Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.

So it makes sense for it to be a bit more paranoid.

There are other possible architectures probably but for now I think nobody uses them. See e.g. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/


Replies

redox99 • today at 1:15 PM

It's not the same thing.

Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.

➕ show 1 reply