logoalt Hacker News

jsw97today at 12:53 PM1 replyview on HN

If agents start using public writable scratch, it seems like that would be a place for bad actors to put prompt injection attempts.

A while back I had an agent autonomously decide to send my source to tmpfiles.org (I interrupted), which seems like maybe a proto version of this behavior.


Replies

pixl97today at 3:22 PM

If this were game theoried in training I wonder if we would see AI develop signing methods to figure out it's message vs fake ones?