> Without training cost you can infer only the marginal cost of serving this kind of models.
Which is by far the most interesting number of the two.
> Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)
If you get close in output quality, then does that matter?
> Which is by far the most interesting number of the two.
Only if you don't have to continuously train new models, and you are not at a runway risk.
> If you get close in output quality, then does that matter?
When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?