logoalt Hacker News

pixelpoetyesterday at 8:34 PM1 replyview on HN

Zen3 isn't the best performing CPU today, though of course it'll never approach GPU level (which is pretty much ASIC level for matrix muls with special units). CPUs are getting dedicated "AI accelerators" too so it'd be interesting to compare per watt. The real limit is almost certainly memory bandwidth, not flops.

It would also be very interesting to see someone like Fabien Giesen / ryg do a maxed out AVX512 version for Zen5. His code's so fast it makes Intel 13900k's self destruct.


Replies

zzzoomyesterday at 9:28 PM

Matrix multiplication is one of the few operations that isn't regularly limited by memory bandwidth. BLAS implementations come with several heavily optimized, architecture-specific versions of sgemm.

show 1 reply