logoalt Hacker News

danielmarkbruceyesterday at 10:39 PM1 replyview on HN

Respectfully, go build one, including doing RLHF and RLVR. Those phases generate lots of tokens, then get scored on the entirety of the output, then optimize based on a scoring of that output. It doesn't check a "prediction" against what was actually "next" in data, because there isn't any "next token" data it's training on.


Replies

angoragoatsyesterday at 11:06 PM

> It doesn't check a "prediction" against what was actually "next" in data

Literally no one here is claiming that it does. This is one of the many flaws in the article.

show 1 reply