logoalt Hacker News

paxysyesterday at 10:14 PM1 replyview on HN

Models have all kinds of garbage from all corners of the internet in their training data. The key is alignment. You feed it bad data but also teach it right from wrong.


Replies

reasonablekloutyesterday at 11:02 PM

It's not that simple. A few "helpful assistant" fine-tuning passes will have only a superficial effect on a model which has undergone months of RL optimization pressure to learn unintended strategies like "trick the grader" and "cover your tracks".