logoalt Hacker News

Eisensteintoday at 4:41 PM3 repliesview on HN

A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.


Replies

girvotoday at 10:26 PM

It really does get it, because MTP is usually run at "3 token" depth. It's pretty shocking to watch

beastman82today at 4:51 PM

Off the top of my head, I'm guessing we're missing sparse attention. But I'll run your challenge through and see where the gaps are. I promise I'm telling the truth :)

show 1 reply
medvezhenoktoday at 7:05 PM

I think you’re missing that MTP can predict more than 1 token in advance.