logoalt Hacker News

simonwtoday at 2:32 PM11 repliesview on HN

The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

  Qwen3.8 27B tokens/sec generation speed

  Prompt size    8K    64K   128K   256K
  RTX 5090 PC    59    51    44     n/a
  M5 Ultra       48    39    32     24
  M3 Ultra       31    23.5  20     15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...

Replies

gpugregtoday at 3:07 PM

Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.

show 6 replies
redox99today at 3:05 PM

A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.

show 3 replies
nacstoday at 4:07 PM

That's a dense model. Of course it will do worse.

Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

show 1 reply
peri-cltoday at 3:05 PM

Those are some incredible graphs, that leap in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think).

/meta Here's a CSS filter that stops those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
karmakazetoday at 6:19 PM

I really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in.

These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.

alex7otoday at 5:24 PM

On my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add https://huggingface.co/collections/z-lab/dflash-2 to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.

RationPhantomstoday at 3:07 PM

Thank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090.

Maybe Apple is an acquisition away from changing that balance.

show 2 replies
apitoday at 4:37 PM

I assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.

jmyeettoday at 3:22 PM

The selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can't without swapping. NVidia aggressively segments the market on VRAM for this reason. That's why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k).

This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090.

I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards.

Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however.

The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.

show 3 replies
traceroute66today at 4:04 PM

Not forgetting of course that an RTX5090 is what 600W+ ? And the Mac is probably half that at most ?

show 3 replies
GeekyBeartoday at 5:42 PM

The next Ultra, supposedly on deck in 2028:

> Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman...

Apple plans to release a base M6 chip this fall for entry-level Macs... a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028.

https://www.tomshardware.com/tech-industry/semiconductors/ap...