logoalt Hacker News

GodelNumberingtoday at 4:41 PM4 repliesview on HN

Back of the envelope calculation (could be off, correct me if I am)

If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.

Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!

As long as a company believes that the openweight models will continue to get more capable and 'AI is here to stay', this model provides the first solid footing for a decision to just buy a rack.


Replies

FuriouslyAdrifttoday at 6:36 PM

We're running Kimi 2.8 on a $107k server and getting around 50k tokens per second or better on most things.

We've already saved money compared to last years token cost on Claude/Gemini

show 4 replies
999900000999today at 4:52 PM

And hire 2 or 3 dev ops to keep it running ?

That another 400 to 700k.

It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

show 9 replies
kcbtoday at 6:20 PM

You also need a place to put it. With liquid cooling and extremely dense power capability. Your typical colo or closet server room isn't going to cut it.

recklesstoday at 4:52 PM

I think the licensing that would likely apply to a company that's able to afford ~$6M rack and the associated infrastructure muddies this somewhat

show 2 replies