logoalt Hacker News

gpmlast Saturday at 7:40 PM1 replyview on HN

I strongly suspect the flip side is that in the future it enables you to train smarter models by "distilling" the end result of the super duper heavily thinking models.


Replies

esafaklast Saturday at 8:10 PM

But these models already distill the smarter American ones ;)

show 1 reply