logoalt Hacker News

Slartietoday at 7:23 AM1 replyview on HN

I'd guess the slowness is mostly due to there currently being only one provider, Moonshot AI. And they are overwhelmed with demand.

Let's judge the speed of the model when its weights are released and every inference provider on the planet offers it, so demand can spread out a bit.

It's the same topic with token budget comparisons and subscription pricing - don't people understand that this doesn't really matter for open weights models? The pricing is going to be determined by the inference providers, and until they had a chance to evaluate the model on their infra and set token prices accordingly, one doesn't really have anything tangible to compare with other open models nor with closed ones.


Replies

epolanskitoday at 7:29 AM

Even if hardware capacity increases, it seems clear it uses way more tokens, so I don't expect parity with other competitors on that front.

On the other hand I expect K3 future refinements to be massive and more efficient.

show 1 reply