I'm sure that local LLM will be far cheaper
Depends on what level of intelligence you're wanting to use. A vanishingly small number of people can or would want to go to the hardware expense of running something like GLM 5.3 Flash, much less something like K3.
And if you want Astra/Fable/Opus frontier level, then there's no option at all.
But if you don't need that, or you don't need speed... That opens up the discussion. I've been impressed even with how Siri's been doing with the Apple Foundation Models in MacOS/iOS 27 given how small they are.
Edit: I can't even fully spec the M5 Ultra Mac Studio you'd need for GLM5.3 Flash since 512GB isn't available yet, but it's already at $9500 for 256GB RAM.
Best comparison that has occurred to me is the cost a loaf of bread's ingredients might be slightly cheaper than a baked loaf, depending on how you source it. At home you get total control and know what's going in to it. Yet bake at home is still a niche, perhaps a hobby. So I say as someone who's spent hundreds of hours tinkering with local inference, go for it for anyone reading. But most people just want ... some slices of bread, you know?
Wouldn't be so sure, at least not for a GPU-based system. Quick math for a 5090 running at around 500W generating 100 tokens per second (reasonable for Qwen 3.8-27B) is around 2-3 kWh for a million tokens, which is around $0.60 for some mix of off/on peak electricity rates.
This is in the same ballpark for that same model on openrouter (https://openrouter.ai/qwen/qwen3.8-27b), highly dependent on input/output mix. Deepseek is a much more capable model that you can't run locally on normal hardware, and their rates are insanely cheap ($0.04 / $1 per 1M).
And so far I haven't considered the cost of the hardware. I happen to have a gaming PC that can be put to use on inference when not gaming, but given these numbers I don't think I would buy new hardware to do inference at home. Unless my math is wrong, it seems you're way better off paying for Deepseek than running Qwen or some other locally runnable model yourself. Of course if you have specific privacy requirements or prefer something unique about a particular model you can run locally, the equation changes.