logoalt Hacker News

CoolestBeansyesterday at 8:29 PM0 repliesview on HN

It is probably why RLHF was the secret sauce to make LLMs vastly more useful. Obviously it makes their output more likely to be aligned. But also by mimicking conversation it makes you prompt it better. It manipulates you into not only providing a prompt that will be more likely to give a response that works but also reveals more information because you've spent all your life talking to other people. And of course more information in means it will be able to find that intersection of information you care about.