logoalt Hacker News

whatshisfaceyesterday at 7:55 PM1 replyview on HN

The method of the NS advance involved RLHE (reinforcement learning via human example), and that is only open-ended if users continue to advance the frontier within chats ahead of publications.


Replies

thorumyesterday at 8:32 PM

Sure, but the point is that the labs use more powerful internal models for research work, not public models. Public models tend to lag the internal frontier by a decent margin, and are constrained in other ways by monitoring. It’s just not a useful indicator.