logoalt Hacker News

vorticalboxtoday at 6:20 PM1 replyview on HN

Putting the whole model in memory is far faster then swapping to disk.


Replies

giancarlostorotoday at 6:46 PM

For local inference the cost of "speed" is not that bad I would think? I wouldn't mind a bit of a delay if it means I can run much larger models on my Mac.

show 1 reply