My biggest problem with running local LLMs on my M4 Max/128GB RAM is the prefill latency.
I've since acquired two DGX Sparks, and it feels so much snappier.
m5 max really fixed pp with the better matmul support, im sure the m5 ultra will be even crazier
the sparks have much slower memory bandwidth is the trade off
Would you mind sharing your local Mac setup and which models you currently use and whether it’s GGUF or MLX? I’ve the hardware same specs.