logoalt Hacker News

hyperpapetoday at 8:10 PM1 replyview on HN

Maybe, but that's sort of begging the question that those open weight models aren't significantly trained using "distillation"[0]

[0] not technically distillation. https://thomasdullien.github.io/posts/2026-06-15-rl-economic...


Replies

verdvermtoday at 8:25 PM

distillation is a minor piece of training data, you have to have a good foundation for it to be helpful, and even if you have good traces, you need a good RL reward scheme at the point it is used (very challenging)