yes, I average 80-120 tok/s on my RTX 3080 with gemma 4 and faster with Qwen 3.5. The main use-case here is just code-monkey agents. I'm not looking for architectural guidance, but an agent to take a spec and complete it.
And is this using conventional model loading (all in VRAM), or are you streaming it in some way?
And is this using conventional model loading (all in VRAM), or are you streaming it in some way?