logoalt Hacker News

tommicatoday at 5:13 AM5 repliesview on HN

What on earth hardwares do you guys have to be able to run 100gb models locally?! That's crazy! I'm here struggling to even get 27b models to run in somewhat usable way


Replies

numpad0today at 9:19 AM

Yeah, here I am sitting deeply deeply deeply regretting not buying couple CMP 170HX at $200 or $350, knowing I could just flip them ethically at purchase price if nothing came of it... I could have just casually built a 128GB dual A100 local AI monster

nozzlegeartoday at 5:24 AM

Haha I'm on an Mac Studio with an M1 Ultra, 64gb ram. I bought it when it first came out, it just happens to be good for local LLMs. I have to use a smaller quant of Laguna S though (I think 4-bit? Not at my machine to check), as 8-bit and full size definitely don't fit in the 64gb I have.

show 2 replies
szniotoday at 7:25 AM

quantized + offload

I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s

show 1 reply
sarjanntoday at 7:16 AM

dgx spark, nvfp4 so I have spare room for KV cache (context)

colordropstoday at 6:37 AM

MoE models can use system memory along with a GPU.