logoalt Hacker News

Foobar8568yesterday at 6:54 PM1 replyview on HN

Memory used : 38GB, and I haven't even started a LLM nor podman, I always fight with memory when using LLM on my mac with 48gb.

And I don't remember to have been able to have pushed to 200k context Qwen 3.6. 3.8 is running on my RTX 5090.


Replies

redox99yesterday at 7:59 PM

Qwen 27B runs very comfortably on a 5090. You need to use Q4 quants and Q8 KV cache. Here's the math

https://news.ycombinator.com/item?id=49514141