I am not a proper developer and only use AI for faster research of topics so please forgive my ignorance. Could one not save a lot of money on tokens by using the 80/20 or 90/10 rule in that 90% of AI usage is on local models and save that last 10% or less for the frontier models where the local model did not meet the needs? Did they cover this and I misunderstood?
This is like the "half of my marketing spend is wasted" quote. The complexity is finding out which half.
To run a local AI that is half decent at research at usable speeds requires hardware that costs thousands of dollars.
Spending thousands of dollars on hardware to save dollars per month on tokens does not make financial sense.
If you run the numbers you'll probably find that using cheaper cloud models makes more financial sense than running those same models locally.
That can work, especially for privacy or repetitive tasks. My experiment focused on subscriptions I already had, but local models are a natural addition to the routing layer.
Unlees you invest thousands your local ai won't even come close to the cheap cloud llms. Local is only worth it if you care about privacy or have a legitimate usage for the hardware otherwise. Money wise it isn't worth it