logoalt Hacker News

fnordpigletyesterday at 11:55 PM1 replyview on HN

When I say distill I also mean mine it architecturally for insights but I doubt seriously the model training is entirely distillation of Claude, it’s almost certainly a mixture of both original corpus and reinforcement as well as distillation. I think it’s a little condescending to imply that these new open models are cheap ripoffs with nothing original to them. These teams and labs are top tier as well, working under unreasonable constraints imposed by the USG. That’s a powerful combination for creativity.


Replies

linkregistertoday at 12:53 AM

I think your characterization of my post, which credits Moonshot's innovation, goes a bit too far.