- Tiny context size or hours to load it
- Hard to benefit from thinking and preserve thinking given token cost.
- Low quants reduce accuracy heavMTP draft can make it make the same mistakes all the time when calling tools, formatting output or following basic guidelines. Otherwise 2x-4x slower.
- K/V quants probably quantized too make things less accurate.
Useful would be combinations with:
- Full context size so it can code and think a bit.
- Draft MTP <= 2 so it doesn't trip
- Q4 quants or better so its accurate
- q8 cache or better so it stays accurate.
- 20 token/s so it finishes while reviewing previous step.
- 1000 tokens/s context load so compactions don't waste 10+ minutes.
- And enough left RAM for 50+ context checkpoints so that it can progress quuckly.
Closest you have is Qwen3.6-35B-A3B-MTP.
Latest gens (Qwen3.8 and co.) are just too big for low specs. 27B dense models seem to be ok for integrated >=92 GiB RAM.
Source: I have low specs and tried them all for agentic use + coding.
I’m using qwen3.8 quants (q3) effectively on a 5070 ti. A lot of it is about guardrails, but you also have to figure out how far (and in what ways) you can push a given model.
27b runs perfectly fine on 2x 24GB at ~100 t/s (4090) with speculative decoding on 8 bit quants
What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
> Flash attention with who knows how much draft ensures it makes the same mistakes all the time and cannot call tools, format output or follow basic guidelines reliably.
> Draft MTP <= 2 so it doesn't trip
I am not sure you understand what either of these things do.
Do you think that FA or MTP are lossy?
Can you please make your substantive points without snark or swipes? This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.
There's a lot of good information here but the comment spoils itself by coming across as aggressive in this way.
This is particularly a problem when responding to someone else's work. We need commenters to point out problems respectfully, not put down what other people have been making.
Everyone's definition of usable is different, but I disagree with your estimation of the specs required to be useful. I am able to do useful coding on Qwen3.8-27b, with 100k context and 16gb VRAM, q8 cache. I feel I'm living right on the cusp... my GPU is old (2016, pascal), so to get usable speeds I have to drop to a Q2 quant - which still gets stuff done, but the difference with q4 is noticeable. Q3 is close enough I don't really notice the difference between it and Q4, but it's too slow on my system. More context would be nice, but it's not that hard to work within ~100k.