logoalt Hacker News

xpcttoday at 2:45 PM1 replyview on HN

It's probably fuzzily fixable by including instruction authority levels in the training data. Can't expect much more than that, given that the model itself is fuzzy.


Replies

QuercusMaxtoday at 8:40 PM

We've already seen that it's possible to trick models into seeing user input as their own "thinking" if you make it sound like what the model writes. While it may appear that it's looking at the tags on the input, in practice that's not as strong a guarantee as you'd hope.

show 2 replies