The page says 170G/s memory bandwidth for the NPU and 1.2T/s for the GPU. Why the discrepancy if it's all "unified memory"? The former is nothing to write home about as far as AI compute is. The latter is really nice.
Which one is it you can run local models on? I suppose the NPU only.
Unified memory is about address space. The bandwidth is still determined by bottlenecks to the processor. CPU/RAM links are still fairly narrow.
I think you misread, it’s 170gb/s for base M6 model and 1.2tb/s for M5 ultra.