logoalt Hacker News

FuriouslyAdriftyesterday at 7:12 PM0 repliesview on HN

We run Gemini fast for random end user queries for general staff.

We have our own on-premise inference server (quad MI300A) that runs Kimi 2.8 extremely well and we transitioned all heavy work to it since it's basically instantaneous for the whole team. It's a good enough solution and we will hit break even before the end of the year already.

Not everyone needs frontier models and availability is frequently much more important than a lot of companies realize.