logoalt Hacker News

daemonk • today at 4:36 PM • 1 reply • view on HN

Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.

The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?


Replies

tnspacetime • today at 5:26 PM

I have not studied it properly too. Good that it has both code and note though.