logoalt Hacker News

TurdF3rgusonlast Thursday at 11:39 PM2 repliesview on HN

The cost of those output tokens is not zero or marginal.


Replies

aembletonyesterday at 9:59 AM

You could run the inference locally using the open weight models. Then you'd still get to use the models without sending money overseas

cobbzillayesterday at 1:52 AM

The first token on a new AI rig costs $X (full capex cost) then every token after that costs virtually zero. Over time the cost/token trends towards zero (modulo opex). That said, AI does have higher opex than general SaaS so it can’t get as close to zero.

But that’s kind of a different question, the running of some service. The “product”, the model, is a collection of files. The “manufacturing” required to add another customer is “send them the files” and has ~zero marginal cost.

show 1 reply