logoalt Hacker News

Ilaurensyesterday at 5:41 PM3 repliesview on HN

These token numbers look off for a 5090. For comparison, an rtx 5000 Blackwell SFF with just 470gb/s bandwidth gets me 30 tok/s on gemma4-31B-QAT. Almost 60 tok/s with MTP enabled. A 5090 should get you much more than that!


Replies

jszymborskiyesterday at 6:04 PM

Thanks for flagging, this is on my local rig and it's driving my display too. I'm curious now, will take a closer look. These are the tok/s as reported by LMStudio.

EDIT: Updating llama.ccp gets me 58 tok/s on Gemma 31b

show 1 reply
mft_yesterday at 6:00 PM

Agree - I got 9.7 tok/s on an M1 Max with unsloth's gemma-4-31B-it-qat-UD-Q4_K_XL.