> their return on capex multiple, their CEO said they make a six-fold profit on their compute capex with 10 month recuperation
I am not familiar with chineese model companies as much as I am with US based ones so I don't have much to say beyond that the CEO is incentivced to pump up those numbers.
> By limiting the number of tokens you use per month? per week, per hour? And by limiting the inference time compute dedicated to each turn in each session.
If this was so simple I don't know why GitHub copilot went to token based billing at my company.
> This does not need to be solved, but rather only quantified. Innovation is needed to be able to reasonably bound this variance for a reasonable subset of tasks
It's much better to make a business case for them after finding this bound right? Currently I can't use copilot for anything serious since I cannot predict how many credits one request is going to consume.
Your point about non frontier tasks using less tokens makes sense. As you said, let's see if it holds up
Return on compute capex is tied mostly to gpu lifetimes so I don't think it will be different for the American companies, who also charge much more being closed source.
> If this was so simple...copilot...
Github copilot still has subscriptions. They moved away from request based accounting to token based accounting for the usage limits, as did cursor, and everybody else.