logoalt Hacker News

giancarlostoroyesterday at 6:46 PM1 replyview on HN

For local inference the cost of "speed" is not that bad I would think? I wouldn't mind a bit of a delay if it means I can run much larger models on my Mac.


Replies

pertymcpertyesterday at 6:52 PM

It's pretty painful to have speeds < 30 tok/sec though. Especially if you're used to API providers at higher speeds. It makes any interactive work almost impossible to do efficiently because you have no choice but to context switch after every request.

show 1 reply