logoalt Hacker News

zozbot234 • today at 2:13 PM • 0 replies • view on HN

> you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would.

This has been done quite effectively with open weight "abliterated" models. You figure out under what sorts of circumstances an undesired behavior is elicited (this is all about pure simulated rollouts, no real-world action required) and what's the closest equivalent you would prefer, then surgically take out the unwanted behavior and shift the model towards the preferred one. It's similar to how RLHF works but much more precise in targeting what's unwanted and limiting impact on the rest of the model as a whole.

This is relevant to real-world safety scenarios, e.g. there's been anecdotal evidence that Claude Fable has been "steered" away from active cyber offense (this is very similar to how abliteration works) and will just not do that even if you otherwise manage a "universal" jailbreak of the model.