logoalt Hacker News

HarHarVeryFunnytoday at 3:07 PM1 replyview on HN

It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.


Replies

upbeat_generaltoday at 3:49 PM

This is just not how it works - it is perfectly plausible (in fact the most likely) that pretraining was well finished by the time Fable was released. A bit of extra distillation takes far less compute.

Moreover, it’s very plausible (and expected) to use multiple clusters and GPU types for RL rollouts which could very well not be included in this count.

No part of this pipeline is fixed in stone.

show 2 replies