You know what, considering there was a recent "small" open weights LLM released recently that meets 90% of my coding needs I'm inclined to agree.
Qwen3.8-Flash-Next - relatively small, it runs on 6 6 year old GPUs on my home PC happily running 5 simultaneous 262k sessions with additional 10 cached in RAM (bought back when you didn't have to remortgage your house for Ram) and it has been the first local model that is not a toy.
But there is a class of problems where I still reach for Anthropic's fable...
However, I have a hunch bordering with certainty Anthropic is achieving such great results by doing a lot of harness tricks.
For example opus 4.8, is not much better on coding than before mentioned Qwen model, but gets amazing results on factual knowledge stuff (the knowing all works of Shakespeare thing). How hard would it be to add a general knowledge RAG to requests that contain relevant questions and beat all benchmarks like that? Not very hard.
So I think there is big innovation to be had in harnesses, routers, inference and so on.
As to money spent on AI per developer my current client (a fortune 200 software company) spends $500 per month. That is $6k a year. A lot more than your examples. And many people run out of their quota pretty quickly.
Can you tell more about the class of problems that need Fable?
I was inclined to listen to your take… then you started talking about having 6 graphics cards in your computer and acting like that’s a normal thing that people do…