logoalt Hacker News

cmrdporcupinelast Sunday at 12:09 PM0 repliesview on HN

For whatever reason prefill (on my DGX Spark) is faster with the Gemma models than Qwen 3.6 models of similar size. On vLLM anyways. Likely just deeply tuned code contributed to vLLM by Google?

vLLM gives me ~7000+ tok/sec with Gemma 4's MoE model. Vs ~6000 tok/sec for Qwen 3.6 MoE.