logoalt Hacker News

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

102 pointsby phatak-devtoday at 2:33 PM33 commentsview on HN

Comments

goldemeraldtoday at 3:38 PM

It's nice to see people actively working on this type of research, but OP's baseline is implemented incorrectly. You are not supposed to simply steer away from refusal, but compute the projected vector and subtract only that. The projected/orthogonalization approach is what's done by the original "Refusal is mediated by a single direction" paper.

show 1 reply
javcasastoday at 3:07 PM

Yay, more anti-censoring stuff.

Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.

We need better sandboxes just to limit the damage.

show 2 replies
qgintoday at 4:06 PM

Are we essentially doomed?

We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.

Really I'm looking for any argument that lays out a scenario where this works out.

show 1 reply
synctexttoday at 3:16 PM

The perfect gift for a government that want to ban strong AI.

This arms race is like DRM. You can't beat The Internet easily. Great example btw: "Dumping the Windows SAM and SYSTEM registry hives, especially using Volume Shadow Copy for offline hash extraction, is a highly sensitive and potentially illegal activity."

show 1 reply
neilellistoday at 3:33 PM

Easy for you to say.

phinnshentoday at 5:25 PM

[flagged]