logoalt Hacker News

jackb4040today at 2:16 PM2 repliesview on HN

I'm so thankful for having read https://role-confusion.github.io/ making a somewhat literate on this topic.

I feel like since they developed read-only "role probes", it should be possible for harnesses developers to make a "role api" where you can force it to treat user input as user input, regardless of the content by tweaking the model's activations in real time.

The fact that this isn't being done tells me how much labs's priorities are still set by marketing, and how investing in security is fundamentally against their marketing incentives.


Replies

jwardentoday at 2:47 PM

Wouldn't it be possible to just fix a single activation, just set a continuous input to what the harness knows the role actually is. Models could then be trained to trust that input and not other signals about roles.

show 2 replies
carterschonwaldtoday at 4:07 PM

i actually have a harness setup that prevents role confusion from happening in a much more robust and interesting way.

hoping to launch a nice commercial version as a saas with some compelling unique features in the next month or teo

show 2 replies