logoalt Hacker News

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

87 pointsby marcobambinitoday at 2:12 PM33 commentsview on HN

Comments

pjatoday at 5:39 PM

That README hits all my “this is authored by an LLM” instincts. I presume the codebase is also written by an LLM?

show 8 replies
rounduptoday at 8:06 PM

How does this project compare to https://github.com/gavamedia/deltafin ?

bgirardtoday at 8:05 PM

Approximate calculation is putting the cost at ~$5 per million tokens (assuming 42W sustained, 20¢/kWh), and that's excluding hardware and other costs.

Catloafdevtoday at 7:52 PM

Neat! But, what do you do with a 0.5tk/s LLM?

Have you tried running it via llamacpp or other software that supports naive SSD offloading to compare speeds?

show 2 replies
righthandtoday at 8:22 PM

I couldnt find anything explaining the name of this company on their website but is it okay that they’re riding on the name of an open source tool?

SQLite code itself is public domain but I’m not sure about the name.

SSilver2k2today at 7:41 PM

This sounds a lot like what the colibri project did for GLM-5.2. I'm a fan so keep at it!

justvugg.github.io/colibri

cadamsdotcomtoday at 8:14 PM

Dear creator: you didn't ship the first draft of your code - why did you ship the first draft of your README??

show 1 reply
herftoday at 6:40 PM

So if this Mac uses 30-50W, that's 40-60 tok/Wh...vs maybe 80k for a modern GPU cluster? So that's about 1000-2000x more power for the SSD streaming, unfortunately.

cjbprimetoday at 5:29 PM

Does it not use Metal, on macOS? Would it be faster if it did?

show 2 replies
jpecartoday at 6:04 PM

Where can this 1tb k3.waste be downloaded?

show 1 reply
logicalleetoday at 6:14 PM

Interesting project. The headline number (29 GB of RAM) is for 4k context.

From what I've read elsewhere, Kimi K3 is quite verbose in its thinking. At the quoted rate, it would generate only a total of 1.8k tokens in 1 hour. Is that enough for it to get any thinking done and produce output on more complicated prompts?

show 1 reply