logoalt Hacker News

tjwebbnorfolkyesterday at 11:43 PM0 repliesview on HN

Inefficient per watt, yes, but local inference capacity is greatly underutilized in aggregate. If a model can run on a machine that already exists, that's a bunch of additional chips that don't need to be built.