logoalt Hacker News

kevin42yesterday at 3:38 PM1 replyview on HN

What hardware do you run? I have a first-gen mac studio, and I just run cmake and build with no special options. Same thing with llama-server, I just specify the model and use the built-in web UI.

For reference, I get ~26 tok/sec with the new Muse 30B model.


Replies

dofmyesterday at 3:43 PM

An M1 Max MBP manages roughly 10 tok/sec without the Dflash speculative draft support so that tracks; the M1 Max apparently has trouble actually saturating its memory bandwidth.