Why not train another smaller LLM to give the same answers as Kimi K3?
Exactly my machine 64GB M1 Max So happy about this! ♡
idk how people access (soldout) and even afford 512GB RAM MacStudio's. Isn't it $40k or so?
Would be interesting to see how fast it would be on 4x mac studio 512gb machines.
The title should probably be edited to specify "M1 Max" instead of "M1 Mac". You aren't running K3 on a base M1 anytime soon. Either way, still a very impressive project.
60s/token - if only there was a way to drop that "s" this would be amazing
Anyone who knows the state of NVMe hardware more than me know if this would obliterate the lifespan of your drive? Seems like the biggest limitation to me (some people are probably fine with letting their Macs churn over the weekend).
0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?
Now set it up with an agent and a permanent `/goal` to say it cannot stop until it has solved for speed, then leave it on and livestream so we can all see when it becomes exponential. Could have the Eternal Jukebox playing in the background!
> ~60–76 s/token
I don't know if I'd call this "running"
Cool gimmick
under 0.02 tok/s
[dead]
SSD streaming on an M5 Max 128GB: https://x.com/antirez/status/2082136334160818528
Soon decent speed across two Mac Studios with 512GB of RAM.