It doesn't really matter what the AI companies want, if they misjudge the market somebody will just start a competitor and take it.
The relevant questions are: 1) Where are the economies of scale in the technology stack? 2) What's "good enough" to consumers, and how does that stack up with the relevant computing power needed? 3) What are the transaction costs along various system boundaries?
I think that the biggest force keeping inference in large centralized services is simply that provides a better product for the average user who doesn't care about local control (and the average user doesn't care about local control; indeed, for most people it's a misfeature, as then they have to administer their own hardware). HN is full of nerds that want to own their own stack; they're willing to put up with a little loss of capability to run Qwen 3.8 locally. But from the folks I know that have tried it vs. Claude vs. Codex vs. Antigravity, the local models are still pretty weak compared to what you can get by paying for a service. Most people will just pay for the service until the performance becomes indistinguishable and the price becomes less.