"it'll barely run on an M5 Max "
The max version I could order now with 128 GB?
If so, the price for local inference would be 12 000 € vs 500 000 € for a B300.
There's also the 2x spark way, which should be ~8k eur? Someone down the thread reported ~60tps for 2x sparks. That's totally usable for local inference.
You can also do 2x 6kPRO in a workstation, for ~20k.
500k is for the 8x B300 version. Which is the only one you can buy atm. But technically a B300 card is more like 60k, just impossible to get.
I'm running a useful quantization of the previous version of Deepseek-V4-Flash -- quite well but with so much fan noise -- on a MacBook Pro M5 Max with 128 GB.