There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is lying through censorship: https://arxiv.org/html/2603.05494v2
Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.