logoalt Hacker News

jszymborskitoday at 5:24 PM4 repliesview on HN

On my RTX 5090, Qwen3.6-35B-A3B is my go to for coding, but Gemma-4-26b-a4b (Q4_0) is blazing fast and super high quality. I sometimes use the dense cousins of these models (Qwen3.6-27B QB_0 and Gemma4-31B-QAT Q4_0) but they can be very slow. Integrating web search and code calling via Open-WebUI makes it so that I usually don't need to reach for Claude.

Here are the tok/s I get:

- Gemma-4-26B-A4B (Q4_0) = 214 tok/s

- Gemma4-31B-QAT (Q4_0) = 58 tok/s

- Qwen3.6-35B-A3B (QB_0) = 30 tok/s

- Qwen3.6-27B (QB_0) = 9 tok/s

EDIT: Updated tok/s after updating llama.cpp


Replies

NorwegianDudetoday at 7:04 PM

You can get easily over 100 tok/s on the Gemma 4 31B QAT if you enable MTP. Same goes for the 27B. I'm getting 880 tok/s of throughout on a single RTX 5090 for batch tasks.

Ilaurenstoday at 5:41 PM

These token numbers look off for a 5090. For comparison, an rtx 5000 Blackwell SFF with just 470gb/s bandwidth gets me 30 tok/s on gemma4-31B-QAT. Almost 60 tok/s with MTP enabled. A 5090 should get you much more than that!

show 3 replies
flockonustoday at 6:45 PM

To the models you say for coding, can you give an example of what you code?

I find them insufficient for my projects (mid sized), but curious what people see working.

show 1 reply
colordropstoday at 6:12 PM

I see a lot of conflicting info about whether Qwen3.6-35B-A3B or Qwen3.6-27B is more capable. Is it one of those things where "it depends"?

show 1 reply