logoalt Hacker News

hecturchi • today at 5:27 PM • 6 replies • view on HN

- Tiny context size or hours to load it

- Hard to benefit from thinking and preserve thinking given token cost.

- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.

- K/V quants probably quantized too make things less accurate.

Useful would be combinations with:

- Full context size so it can code and think a bit.

- Draft MTP <= 2 so it doesn't trip

- Q4 quants or better so its accurate

- q8 cache or better so it stays accurate.

- 20 token/s so it finishes while reviewing previous step.

- 1000 tokens/s context load so compactions don't waste 10+ minutes.

- And enough left RAM for 50+ context checkpoints so that it can progress quuckly.

Closest you have is Qwen3.6-35B-A3B-MTP.

Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.

Source: I have low specs and tried them all for agentic use + coding.


Replies

kennywinker • today at 5:55 PM

Everyone's definition of usable is different, but I disagree with your estimation of the specs required to be useful. I am able to do useful coding on Qwen3.8-27b, with 100k context and 16gb VRAM, q8 cache. I feel I'm living right on the cusp... my GPU is old (2016, pascal), so to get usable speeds I have to drop to a Q2 quant - which still gets stuff done, but the difference with q4 is noticeable. Q3 is close enough I don't really notice the difference between it and Q4, but it's too slow on my system. More context would be nice, but it's not that hard to work within ~100k.

ohyes • today at 6:02 PM

I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.

mirekrusin • today at 6:33 PM

27b runs perfectly fine on 2x 24GB at ~100 t/s (4090) with speculative decoding on 8 bit quants

➕ show 1 reply
ranger_danger • today at 5:31 PM

What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).

➕ show 2 replies
qeternity • today at 6:55 PM

> Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.

> Draft MTP <= 2 so it doesn't trip

I am not sure you understand what either of these things do.

Do you think that FA or MTP are lossy?

➕ show 1 reply
dang • today at 6:59 PM

Can you please make your substantive points without snark or swipes? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.

There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.

This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.

➕ show 1 reply