logoalt Hacker News

lnenadtoday at 4:29 PM6 repliesview on HN

I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.


Replies

springtimesuntoday at 5:15 PM

I’ll be very curious what you get with DDR4. I also almost went that way. I have an Epyc DDR 5 rig and the best I see is 10 tok/s. Caveat being that’s at Q8 and a 4090 doing pre fill so it could be pushed up.

The surprising thing for me is how much work you will need to cool the banks if you’re near your memory ceiling. My memory starts soft throttling at about 74C (dies may be hotter, that’s the bank temp) and will turn down speed to try to stay below 80.

Happy to send my llama.cpp config settings if you want it.

show 2 replies
fsutstoday at 8:26 PM

It’s not unified ram? I.e VRAM so it will struggle

show 1 reply
0x457today at 4:42 PM

Depending on which Epyc you got it might be slower than 1/5 of the speed.

show 1 reply
nazgulsenpaitoday at 4:54 PM

Curious about that price, if you don't mind sharing a ballpark

show 1 reply
guybedotoday at 5:49 PM

I have a dual epyc + 1TB RAM. I could push glm 5.2 to 7 tok/s CPU only.

jchwtoday at 4:35 PM

Honestly I suspect neither of them will be performing terribly well but with DDR4 3200 RAM I wonder if you'll be counting tokens per second or seconds per token. I mean, you do at least get a lot of memory channels at least, compared to consumer PCs. I am curious to hear what performance you get, I feel there is not enough information out there on what different setups manage to eek out.

show 2 replies