logoalt Hacker News

ameliustoday at 8:21 AM2 repliesview on HN

> Without training cost you can infer only the marginal cost of serving this kind of models.

Which is by far the most interesting number of the two.

> Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)

If you get close in output quality, then does that matter?


Replies

embedding-shapetoday at 8:43 AM

> If you get close in output quality, then does that matter?

When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?

show 1 reply
vb-8448today at 9:51 AM

> Which is by far the most interesting number of the two.

Only if you don't have to continuously train new models, and you are not at a runway risk.