The perf delta is smaller than I thought it'd be given the memory bandwidth difference. I guess...

crapple8430 • last Thursday at 1:32 PM • 0 replies • view on HN

The perf delta is smaller than I thought it'd be given the memory bandwidth difference. I guess likely comes from the Blackwell having native MXFP4, since GPT-OSS-120b has MXFP4 MOE layers.

The NVLink is definitely a strong point, I missed that detail. For LLM inference specifically it matters fairly little iirc, but for training it might.

alt Hacker News