I was already rolling around the idea of a 128GB M5 Max MBP. Now this!
A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...
IDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs?
Just out of curiosity, why run "local-local" when you could just set up a Mini or Studio at home and query it over http? [edit] whole conversation about this in another thread https://news.ycombinator.com/item?id=49433413
I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out.