The cost of those output tokens is not zero or marginal.
The first token on a new AI rig costs $X (full capex cost) then every token after that costs virtually zero. Over time the cost/token trends towards zero (modulo opex). That said, AI does have higher opex than general SaaS so it can’t get as close to zero.
But that’s kind of a different question, the running of some service. The “product”, the model, is a collection of files. The “manufacturing” required to add another customer is “send them the files” and has ~zero marginal cost.
You could run the inference locally using the open weight models. Then you'd still get to use the models without sending money overseas