1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric.
For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per second prefill and 50+ tokens per second on generation. This is actually quite usable and near parity to cloud.
They mention "adds the GPU Neural Accelerators." which, if exploitable for LLM loads, would probably help the prefill a lot
How much would is the cost for that machine though, I'm pretty sure I could just buy tokens from a provider and never run out of money for 10 years, and get far better quality output because inference is being served by professionals on far better hardware and this machine would be obsolete long before that as well. Hosting local seems like a possibly the dumbest thing you could possibly do from an economics perspective. And don't hit me with the privacy argument because everyone saying they care about privacy uses fucking gmail, whatsapp and instagram all day long.
Yes, and they specifically mention "Up to 10.7x faster LLM prompt processing in LM Studio" which is probably using the neural accelerator for prefill.