logoalt Hacker News

infectotoday at 3:01 PM1 replyview on HN

Genuine question. Is there a time factor part of that equation?


Replies

HarHarVeryFunnytoday at 3:07 PM

It's obviously a question - did they just train for 10x as long due to having 10x fewer GPUs, but then that spoils the claim that they distilled Fable which was only recently introduced. Now doubt they did use some training data generated from older US models though.

show 1 reply