logoalt Hacker News

tripleee • yesterday at 7:03 PM • 1 reply • view on HN

What models are you running locally? Are you banking on them improving or do you think they're good enough today? 32GB of VRAM there wouldn't be close to enough to run the best local models.

I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.


Replies

latentsea • today at 12:38 AM

Qwen3.8-27B is a huge step up from Qwen3.6-27B. That release only happened relatively recently but that felt like the 'Opus 4.5' release turning point that SOTA models experienced back when that came out. It was the first time I felt like local models are actually good enough to use as daily drivers now. So, it was only after that point that I switched.

Qwen3.8-Flash-Next is better still if you can run fast enough. If you have a dual R9700 setup you certainly can. That model is even better.

Qwen4-27B has been announced but not released yet. I'm super pumped for it because I already use 3.8 as my daily driver at home for all my personal stuff, so I'm definitely happy to take an increase in capability.

There is clearly still room for improvement in local models on consumer hardware. With the Qwen 27B models, If you have at least a 5070 Ti I think you can get away with running a small Q4 quant if you use KV cache streaming. The 24GB cards can run Q4 comfortably. If you have a 32B card you can run Q6 comfortably. If you have 48GB ~ 64GB of VRAM you can Q8 comfortably. Using llama.cpp Vulkan let's you pool VRAM across cards (even AMD and NVIDIA etc), so my machine has a 5060 Ti and an R9700.

A dual R9700 rig is really the sweet spot right now with the vLLM-radiance fork. If you can swing a 5070 Ti in there as well to retain some CUDA access, then all the better. That's basically the equivalent to spending 2 years on a subscription, but gets you a system that can run Qwen-3.8-Flash-Next and of course the even more capable Qwen4-Flash when it releases. At the end of the two years it'll run even better models I'm sure.

I'm all in on local now.