Where is that improvement coming from? Hardware is already here to compute gemm as fast as possible.

pdhborges • last Friday at 9:14 PM • 1 reply • view on HN

Replies

leakyfilter • last Friday at 9:17 PM

Raw gemm computation was never the real bottleneck, especially on the newer GPUs. Feeding the matmuls i.e memory bandwidth is where it’s at, especially in the newer GPUs.

alt Hacker News

Replies