logoalt Hacker News

FergusArgylltoday at 6:29 AM1 replyview on HN

More RLVR. Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.


Replies

Gecko4072today at 6:52 AM

Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?

show 2 replies