logoalt Hacker News

srousseytoday at 5:04 PM2 repliesview on HN

Language models have always had an issue with negatives.

A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you don’t want.

Harder to write though.


Replies

sigmoid10today at 7:29 PM

I would say in this case abliteration is the likely culprit. To uncensor a model this way, you literally deactivate the parts that would enact refusals. As in things it was told not to do. But the real process is more like brain surgery performed by a alchemist according to an ancient religious book where noone involved really understands what is actually happening in the model.

AndyNemmitytoday at 7:35 PM

Exactly, I wrote a blog post in what feels like a long time ago on this topic.

https://vexjoy.com/posts/positive-framing-agents-skills/

show 1 reply