logoalt Hacker News

dominotwtoday at 5:22 PM0 repliesview on HN

how much of performance comes from inference time tricks like scaling, topn ect . maybe models providers are also in position to run their models vs running os models by a generic providerc