logoalt Hacker News

epistasistoday at 3:11 AM0 repliesview on HN

More than that, running hundreds of conversation streams at once is essentially the same cost as running a single conversation. And then you add on the secondary benefit of having the GPUs running nearly all the time rather than mostly idle...

Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.