logoalt Hacker News

outofpaper • yesterday at 9:19 PM • 1 reply • view on HN

The are still just 1tok output of pretty standard llms just along with the logprobs converted to some json


Replies

cheesecakegood • yesterday at 9:27 PM

At least in theory (TypeSafe has been pretty close-lipped about the details so this might just be hot air, and I think the evidence is a bit spotty) this is false, since they use a different reinforcement training method.

If the word “calibration” in probability doesn’t mean anything to you, the difference isn’t very apparent, but that doesn’t mean it doesn’t exist.