logoalt Hacker News

KeplerBoy • yesterday at 7:48 PM • 2 replies • view on HN

How does it take fifteen minutes to read <100 GB into GPU memory? Shouldn't that be limited by SSD speed with everything slower than a minute being a terrible ssd?


Replies

boredatoms • yesterday at 9:11 PM

It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp

teaearlgraycold • yesterday at 7:51 PM

A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.