logoalt Hacker News

stri8tedtoday at 5:39 PM1 replyview on HN

This model was likely trained months before deepseek released their paper.


Replies

manquertoday at 6:30 PM

Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain