logoalt Hacker News

A practical guide to running 8x RTX PRO 6000's

21 pointsby Retell15last Monday at 5:48 PM30 commentsview on HN

Comments

srcreightoday at 9:06 PM

Makes you realize how insane the M5 Ultra Mac Studio is. 1.2GB/s bandwidth 512GB memory. Its rated max power draw is just 480W. And it also has amazing M-series CPUs. It costs less than just one of these GPUs which each take 700W to run.

show 2 replies
schaefertoday at 8:21 PM

> We currently have 14x nodes of CG480-S6053 ready to ship.

Oh, okay, so this is an ad.

I do still think it's well written and interesting... But if anything, it's just making me more curious about the newest generation of M5 Ultra. (and less and less interested in PCI-E Gen 5 anything)

show 1 reply
jimmoorestoday at 8:49 PM

These people have zero idea what they're doing. Not a single mention of pipeline parallelism that would actually make the setup useful to run a big model.

show 3 replies
RachelFtoday at 8:25 PM

For those who can't afford RTX 6000's you can unlock around 20% increased card to card speed on consumer GPUs using this library:

https://github.com/aikitoria/open-gpu-kernel-modules

The hardware supports it, but Nvidia disabled it if the driver detects cheaper cards.

proxysnatoday at 9:19 PM

Utterly awful article. R1? Llama3.1? Not being able to serve larger llms on 8(!) RTX pro’s? You can literally run open weight SOTA models with relative ease. Even 4 GPUs get you there with a bit of elbow grease and compression. Pure slop.

kmike84today at 9:10 PM

Pass. When articles keep mentioning models like DeepSeek R1, or Llama 3.1, or Qwen3 32B, it is a pretty robust indicator of AI slop. LLMs love to suggest DeepSeek R1, etc. - training data cut-off?

No person with real practical experience and real use cases will be using these ancient models as examples, when talking about local LLMs.

xysttoday at 9:00 PM

SLI is relevant again

CamperBob2today at 8:47 PM

4x RTX 6000 Blackwell cards is a good place to be if you can't swing 8 of them, or if you don't have the power or cooling to run that many. A system based on 4x RTX6K can run GLM 5.3 at NVFP4 precision [1] from a US-standard 120V 20A circuit when derated to 300W, and give you a better pelican than Fable 5.1 [2]. What's not to like?

(Edit: I'm mistaken here, the pelican didn't come from Flash on 4 cards but from the full GLM 5.3 model on 8. But the Flash model is still crazy good for its size.)

1: https://huggingface.co/local-inference-lab/GLM-5.3-NVFP4

2: https://crimson-jeri-74.tiiny.site/

show 1 reply
robotnikmantoday at 8:44 PM

Now if only I could afford 8 RTX PRO 6000's

show 1 reply
christkvtoday at 8:23 PM

600W * 8 just for the GPUs when maxed out (besides the cost). Def nothing for my home lab.

show 1 reply
varispeedtoday at 8:20 PM

> (~9.2 million tokens node-wide at 4k context).

stopped reading after that. What 4k context would be usable for?

show 2 replies
inventor7777today at 8:54 PM

This will really help when my 8 RTX PRO 6000s ship. /s