Sorry for being lazy, but is there a rough breakdown like "You get sonnet level for M5 and Opus for M5 pro, etc.", or is it still speculative. Or put simpler, do you get Opus level for the 256GB M5 Max?
Opus is probably ~2T parameter model, so that would probably not run on these. More like Sonnet.
[dead]
For local LLMs with a Mac, rule of thumb is you always want an Ultra (due to memory bandwidth). Even an M1 Ultra is superior to an M6 Pro in this regard.
There are no configurations even close to running something comparable to frontier model variants, they're simply far too large, but something like full precision Qwen 35b or DeepSeek 70b at 50+ t/s is well within available configuration, and potential for plenty of room for large context sizes.