logoalt Hacker News

sharmajaiyesterday at 8:34 PM2 repliesview on HN

I am getting 14 t/s on my 16 GB card at full context with the UD-Q3_K_XL quant. Model link: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF.


Replies

zenopraxtoday at 1:44 AM

For some reason the unsloth models leave hardly any room for context. I've switched to the regular (non-unsloth) and get about 25 t/s and get about 80,000 more context tokens for the same quant.

Balinaresyesterday at 9:53 PM

Wow, interesting. What KV cache quantization do you use?