logoalt Hacker News

Eridrus • yesterday at 6:14 PM • 3 replies • view on HN

It's actually existing flat per-token pricing that is weird.

Neither encode nor decode are linear in compute, so providers need to price for average expected length.

This is just getting closer to the true cost of generating tokens.


Replies

foota • yesterday at 7:37 PM

My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.

hgoel • yesterday at 8:51 PM

Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.

sebzim4500 • yesterday at 7:21 PM

Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO