logoalt Hacker News

throwaway219450today at 4:55 PM0 repliesview on HN

There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is lying through censorship: https://arxiv.org/html/2603.05494v2

Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.