There’s going to be a lot of competition around this model. Let’s see how low AI providers are willing to push prices.
As long as they are transparent about what quant they serve the model and any other optimization they do that also affects performance of inferred tokens.
I think the results might be underwhelming - AI providers need to turn a profit and can't subsidize, and they're working off of the commodity hardware everyone does.
I wouldn't be surprised if they started offering potentiall bad quantizations with much reduced capability at lower prices (without telling the users, of course)