logoalt Hacker News

embedding-shapetoday at 9:07 AM2 repliesview on HN

> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090

Without leveraging system RAM and/or SSDs, I don't think you can, or how exactly are you running this, if this is something you are doing today? With CPU/expert offloading you could probably do it with a 5090 + 1TB of RAM or something like that, but absolutely not on a single 5090 entirely within VRAM.


Replies

robotswantdatatoday at 9:44 AM

Yes hybrid approaches are much better than people realise.

There are a lot of optimisations that are not in the public sphere, source working on start up in this space

show 1 reply
rhdunntoday at 9:32 AM

Yes, that's what I was saying w.r.t. expert offloading, i.e. ensuring that the GPU could fit the active parameters not all the parameters.

show 1 reply