> So the (PCI-E) bandwidth strongly affects time to first token
On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd probably load them on startup before even starting to serve requests