logoalt Hacker News

inigyoutoday at 3:55 PM0 repliesview on HN

There's a silver lining - if the model is trained to defend the Chinese government, that means it has that direction in its semantic vectors and by subtracting that direction always, it can be made to attack the Chinese government