logoalt Hacker News

NooneAtAll3yesterday at 10:06 PM2 repliesview on HN

I don't understand the premise in the beginning

how is running servers supposed to be 0 cost, while running ai inferrence isn't?


Replies

cheema33yesterday at 11:53 PM

> how is running servers supposed to be 0 cost, while running ai inferrence isn't?

For a SaaS business, running servers isn't free. But compared to the cost of running GPUs for inference that you are selling, it almost is. The company I work for is a SaaS company. We have a single production server. A couple of QA servers. All hosted on Hetzner. Monthly cost for servers is less than $400. This generates a few million dollars a year in revenue.

If we were in the business of selling inference, our cost of providing the service, for the same amount of revenue would significantly higher.

Even large businesses like Microsoft, Meta, Google have operated with similar margins. Cost of running servers, compared to revenue was very low. But inference changed that, in a dramatic way.

throwawayffffasyesterday at 10:28 PM

A typical server that costs 10k to 30k to own and operate can serve between hundreds and thousands of requests per second of a traditional web application like facebook for 2-4 kW of power, the marginal cost of each request is effectively zero.

A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost.