logoalt Hacker News

daveguy • today at 3:53 PM • 0 replies • view on HN

That's a great point. A training corpus based on written text will be inherently biased toward confident and right. Then the RLHF exacerbates the problem because people respond more positively to confident and too often assume correct when they read a confident response.