logoalt Hacker News

achronoyesterday at 7:40 PM1 replyview on HN

Yes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed.

Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right?

If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.


Replies

achronoyesterday at 8:15 PM

To whoever downvoted this, it would be helpful if you actually reply with something substantive. The post I'm responding to makes sweeping characterizations, and I challenge it with a relatively good heuristic and indirect evidence, and only get downvoted?

show 1 reply