logoalt Hacker News

fotcorntoday at 4:52 PM0 repliesview on HN

I am also running some old AMD datacenter cards, 2x MI25 in my case. Getting around 30 tokens/second with short context.

Tensor parallel in llama.cpp using RCCL (disabled by default in llama.cpp for some reason). Surprisingly, for these cards HIP is actually faster than Vulkan, unlike the 9070 XT where Vulkan still wins.

ROCm nightlies do actually support these old cards, just not the ROCm stable releases.