You can now, as of this week, run locally models with the same performance of a SOTA model of March this year:
https://news.ycombinator.com/item?id=49214008 https://news.ycombinator.com/item?id=49229621
The FAFO day of Anthropic and OpenAI arrived.
Not many people have $10k of specific hardware lying around
You will spend thousands of dollars on hardware to run those at lower quality (quantized) than the benchmarks where they match March SOTA performance.
Source: I have the hardware to run those locally. I would never recommend it to anyone trying to save money. It’s so much cheaper to pay even Anthropic or OpenAI.