Generation is basically just memory bandwidth math. Each token has to read all the active weights....

phamilton • today at 12:19 AM • 1 reply • view on HN

Generation is basically just memory bandwidth math.

Each token has to read all the active weights. I think that's around 40B parameters active. At a 4-bit quant that's 20GB. With 100GB/s (replace with whatever your bandwidth is) and you get 5 tokens per second.

Replies

SlavikCA • today at 4:07 AM

And with MTP (or other speculation techniques) you can ~double that.

alt Hacker News

Replies