logoalt Hacker News

verdvermtoday at 3:16 PM0 repliesview on HN

most new effort in training comes in the late phase with RL techniques

the pretraining (slurping the internet) only goes so far, the new data being used is from human preferences and agent traces (designed and/or distilled)