logoalt Hacker News

smcleodtoday at 2:04 PM1 replyview on HN

That was mainly before the M4 generation when they didn't have matmul instructions.


Replies

jasonjmcgheetoday at 2:12 PM

M5 prefill is much faster than M4.

I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max.

For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max.

For larger dense models, some fraction of that, but similar multiple.

show 1 reply