logoalt Hacker News

lucisferreyesterday at 9:36 PM3 repliesview on HN

I think it is fair to argue that prompts are not a safety layer at all and can't be relied upon for much.

"Make no mistakes"


Replies

dmixyesterday at 9:59 PM

Yes that's been obvious since the beginning. That's why you should always monitor your agents closely. Just like supervised self driving cars, you have to watch the road and do some hand holding.

The tooling around isolation, logging, and real time security/anonomly detection for regular LLM laptop users is very immature right now. I expect that to change soon.

The alternative is extremely locked down models which is what Anthropic seems to want to do.

show 1 reply
paxysyesterday at 10:36 PM

It’s equivalent to having client-side input validation. Yes it can easily be bypassed, but in the vast majority of cases where users aren’t malicious it gets the job done quickly and cheaply.

8noteyesterday at 10:39 PM

it is a heuristic though, and can be measured as such.

my steel yield strength table is similarly not guaranteed to be correct for the piece of steel that I have in front of me.