Language models have always had an issue with negatives.
A negative like do “not” xyz is just not encoded the same as spelling out what you want vs what you don’t want.
Harder to write though.
Exactly, I wrote a blog post in what feels like a long time ago on this topic.
I would say in this case abliteration is the likely culprit. To uncensor a model this way, you literally deactivate the parts that would enact refusals. As in things it was told not to do. But the real process is more like brain surgery performed by a alchemist according to an ancient religious book where noone involved really understands what is actually happening in the model.