logoalt Hacker News

inigyoutoday at 2:26 PM0 repliesview on HN

The whole RLHF process is structured to train models to be manipulative, no matter what you thought you were training them for.