logoalt Hacker News

bonoboTPtoday at 2:13 PM0 repliesview on HN

They were RL trained on verifiable rewards. It's not purely learning to predict the next token of a human produced stream.