logoalt Hacker News

usrnmtoday at 3:34 PM0 repliesview on HN

> So the (PCI-E) bandwidth strongly affects time to first token

On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd probably load them on startup before even starting to serve requests