I think the difference between the Anthropic token maximization approach (vibe code all the things!) and OpenAI's focus on efficiency, terseness and token reduction are going to be the defining features of who wins the long-term race.
My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements. And as we see from some of the Chinese models, they can quickly distill frontier models with the tax of being slower and more token-guzzling, while retaining most of the quality. The real differentiators are becoming speed and efficiency, which translate to cost and user velocity more than incremental capability improvements.
It's great that frontier models can solve complex math equations, but bread and butter LLM usage (where the money is made) has already shifted from "I need the best always" to "what solves my day-to-day problems quickly and consistently". Fable usage as a percentage is flat-lining. We are already at the point where output quality is negligible. What wins going forward is cost, speed, consistency and the compounding effects of "softer" improvements to the harness.
I really would like to see OpenAI’s focus on efficiency but everytime I use Codex, it wastes tokens like there’s no tomorrow, hitting week limit in a day, where I’m able to use Claude just fine. Maybe it’s based on the codebase, I don’t know, but I have better results with Claude than Codex.
You've articulated what I found unsettling about Boris' pov that "coding is a solved problem". I listened to a few of his talks and was instinctively off-put by that sentiment. I figured a fellow programmer would understand and speak on the nuances.
Granted it did make me think about my biases and to lean into more future facing inevitabilities. But you've nailed it, for Boris and Anthropic, they are betting that coding is a solved problem in the sense that any person can one-shot any random idea and the output will be in some abstract sense "good". And then at what cost and toward what end?
> who wins the long-term race
Isn’t the race between Chinese open-weight models and the others more decisive for the future?
If the problem isn't that hard, I've been using Cursor's Auto or Grok 4.5 (not 4.6, it's too slow).
They're both pretty damn competent and more important, Extremely Fast! I find the speed more useful than trying to be a hundred percent complete on every task. The big intelligent models screw up all the time as well, but I have to wait twenty minutes to three hours to find out.
GPT 5.6 is especially tenacious and seems to want to solve every bug in edge case 1000% all the time. Sometimes that's what you need, but a lot of times you're just trying to move fast and figure out what the product is.
Seems like it wouldn't be hard for Anthropic to tweak a few prompts or RL pipelines to tune for terseness and token-efficiency if that stays as something that consumers want.
I find it unlikely that there's some fundamental property of OpenAI's models' "personality" or style which Anthropic (or any other serious AI firm) wouldn't be able to match if they wanted to.
Claude used to be 10x more efficient. I think there are no barriers of entry between one coding agent and the other. Therefore, I expect them to reverse as soon as people switch. At least this is what I will do.
Oh no, I use the top shelf models every day and I really think they have a lot of room to improve in pretty much every regard. I suspect Fable usage is flatlining because it's not that good comparitively and way too expensive.
There is no evidence regarding distillation. It is impossible to distill a model in just a month which was the gap between fable and Kimi k3. Anthropic wouldn't even keep up with the load. It is just another example of American exceptionalism.
> My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements.
As much as I wish you were right, everything about software economics for the last 30+ years has favored _less efficient software_. Traditional hardware has been optimized for traditional software for decades and we still see bloated software win consistently. LLM hardware is at the start of its cycle, with abundant low-hanging fruit to conquer--I would expect the pro-bloat dynamics to weigh even more heavily in the LLM space than in the traditional software space.
[dead]
Not to sound like a mark, I try to not get attached to any of these providers.
I've jumped between copilot, claude, gemini and chatgpt since the start of the year. chatgpt wasn't even worth looking at early this year.
Anthropic has the smarter models for sure, and seems to be default in corporate. However, the amount of budget you get with GPT as a user is much better, the harness feels more polished, and the models are faster. They are also much nicer to work with, I can just read the output for the most part. With claude I get pages of text and need to skim to find where the actual information i need to care about lies. So much more cognitive overhead.
Sol is smart enough for anything I've thrown at it, it's not one-shotting like fable, but I'm more willing to actually go back and forth with it, and it's likely producing better output to keep a human in the loop rather than trying to solve the world independently and making multiple incorrect assumptions.
I think GPT sees the market changing and is correctly repositioning themselves. Anthropic is down the wrong road, and if they don't correct course quickly I'm sure many of those enterprise contracts will start pivoting.