logoalt Hacker News

joey64yesterday at 4:03 PM4 repliesview on HN

So, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8?


Replies

vanillaxyesterday at 4:11 PM

you cant. The best you can do is Qwen 3.6 27b with a 24gig ( or cumaltive gpus ) to get to 24gb vram. ala 3090, mac with 36gb ram, amd cards, halo strix amd, dgx spark etc. Lots of youtube videos out there.

svachalekyesterday at 7:10 PM

I haven't seen either of these running outside their creator's services yet, but typically you can watch services like openrouter or nano-gpt for it to show up at a (usually small) discount.

Alpha3031yesterday at 4:17 PM

Well, if you're happy with around (as in within an order of magnitude or two of) 0.1 tokens per second... I believe that's around what people are getting when loading MoE weights from NVMe.

show 1 reply
drnick1yesterday at 6:03 PM

There will be smaller versions in the 10-30B parameters range that can run on consumer GPUs.