logoalt Hacker News

StilesCrisis • today at 1:06 PM • 1 reply • view on HN

LLMs are famously bad at determining "if it's not sure it is correct." They are always confident, because a confident tone ranks better in RL.


Replies

wxnx • today at 1:14 PM

> They are always confident, because a confident tone ranks better in RL.

This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).

I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.

➕ show 2 replies