logoalt Hacker News

QuadmasterXLIItoday at 11:40 AM1 replyview on HN

Given an open weights model trained to sometimes bite kids, we can’t train it to not bite kids, even though billions of dollars of research have been thrown at this open problem.

Given an open weights model trained to never bite kids, you can get it to bite kids with 10 prompts and a linear projection, the known simple algorithm doesn’t even need a backwards pass.

yay asymmetry!


Replies

kzrdudetoday at 11:47 AM

Any pointers to more info about that? Sounds interesting.