logoalt Hacker News

What Is RLCD? The Secret Behind Jev

55 points • by tnspacetime • today at 12:21 PM • 7 comments • view on HN

Comments

firejake308 • today at 5:08 PM

> The operational signal was always relative preference. The scalar merely hid it.

Is this another Claude-ism? "X was always Y. The Z merely hid it." Or am I overcalling it?

➕ show 1 reply
WalterGR • today at 3:54 PM

RLCD, not defined in the article, is Reinforcement Learning for Calibrated Decisions.

➕ show 1 reply
daemonk • today at 4:36 PM

Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.

The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?

➕ show 1 reply
jackb4040 • today at 4:21 PM

[flagged]