I kept hearing "just buy a Mac and run models locally, it pays for itself" and wanted to check. Sunk Cost takes a machine, a model and how many tokens you use a day, and works out how long the hardware takes to pay back against renting the same model by the token.
Obviously there are other reasons to buy your own hardware aside from just saving money on llms but this is just looking at it from a raw cost saving perspective.
If you have any ideas on how I can make this more helpful lmk!
Claude Code is $100+ or else be constantly throttled. My usage on GHCP was gonna be $300+ a month.
I paid $1350 and threw an R9700 in an existing machine. That's a 4 month pay off or so.
Plus, I can feed it sensitive data all day and not be worried where it's going.
Not a fair comparison really. If you can run a model locally then you can somewhat train out the guardrails, censorship, and brand-safety. That has value a subscription does not.
Idk about the quality of this setup but just pasting it here as an example. https://explainx.ai/blog/heretic-llm-abliteration-guide-2026
43 years to break even on Qwen 3.8 at 25% the speed of the API, lol. I like the idea of local models for really small tasks like automation/toolcalling, but it will probably never make sense for coding. I tried them and it was just excruciating compared to what you get for $100 a month from a subscription.
I want the autonomy but local models of the size I would have the means to host wouldn't be capable enough. What usecases tend to suit these smaller models that tend to produce incorrect or otherwise flawed responses often? Could they work for anomaly detection and what would a rough architecture look like?
The idea that you need a new machine is pretty ridiculous. I bought a used HP Omen with a 3090 last month for $2k. 57t/s with Qwen 3.8.
I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
The math is wrong, the tok/s is at least 2x that, at least with MTP and Q8 KV which you should always use. And the default tokens a day is ridiculously low at least for coding.
Having said that, it will never pay for itself. A simpler more absolute math is, if I buy a Mac and use it to sell tokens on OpenRouter, will I make a profit? And the answer is no.
This tells me that the max throughput for the models I'm running on my hardware is lower than it actually is. Please allow us to tweak all the variables instead of locking me in to whatever rate you found by searching
Fun feature: can you show some sort of list of the best combos? Eg shortest payoff time for best capability in various situations.
In the “The small print that isn't small” you describe all the disadvantages of running your models locally, but none of the advantages (just check the rest of the comment section for inspiration on that).
Local LLMs are not really about saving money, they're about autonomy. Choose the exact model you want, fine-tune it if you want, and no one can take it away from you.
I have a home server running vibed applications. VPS host would cost $25/mo or $300/yr.
Mac mini can also build iOS applications. I think if you’re a mobile dev, you can have concurrent builds for your agents instead of everyone waiting on a single machine to finish.
Also, i also use my gpu for rendering and learning and playing games.
I wish you could put different setups on here. I have a couple of A6000s on an AM5.
I like this calculator but it’s really wrong at least for dgx spark. I have one and I get 4x the tokens/s .
Yeah no it does not pay for itself just comparing to cloud. Not at these prices at least, people far richer than you or I buy these things wholesale, no scalper, bought a significant amount at cheaper prices, and are wired up the ass with VC money.
The premium is not having your million dollar prize and career stolen by billionaires.
It pays for itself very quickly if you do 24/7 generation. Use an AI agent that orchestrates other agents working on many things at once constantly. If speed is a factor, you'd not buy a Macbook, you'd buy dual RTX 3090s. About the same price, but at least 6x faster than M5 Max. The benefit of constant generation is you can do a lot more research, coding sub-agents, experiments, etc in parallel when you're not "at work". You end up getting a lot more work done than if you only sit there babysitting sessions.
can you add RTX cards too please? 5090 and 6000
Besides from privacy: I already making twice now.you own the hardware and the price had doubled since i bought. Almost tripled. You missed the opportunity and i have 4 of those awesome machines. Cry on.
I sell those to business who need local air gapped requirments and I make a lot more money!
I can run the alliterated models where none of the service prvoider even dare to provide.
THose benefits outweights a few K.
And show me an api provider that allows me to run 10x agents concurrently for 5 days straights .
I was just gonna throw a beefy Ryzen into an ATX chassis. I don't want to pay Mac prices
should throw in a tt-quietbox
[dead]
[dead]
It pays off instantly, because OpenAI/Anthropic can no longer see what I'm doing and that's worth a lot of money to me. If I am offloading some of my thought processes to a machine, I want to own that machine. And if I finetune the model, I can gain access to parts of thought space that are cordoned off by OpenAI/Anthropic/Alibaba/whomever due to their "alignment" efforts (i.e. alignment to the AI company rather than me). Otherwise, it's like if someone else owns a part of my mind and has a backdoor into my mind.