One of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything.
So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it.
So the competitive pressure and predictability offered by open models is helpful for users
> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.
FlashAttention was a hell of a drug.
> if you really want Kimi K2 instead of K3 you can still use it.
I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.
Quantization also started picking up around then, as well as distillation into smaller models
Did they actually ever cut the price on gpt 4? The oldest versions of it in the api still seem stupidly expensive? There were definitely price cuts as they introduced the turbo models and stuff, and new versions of each model might have gotten pricey cuts, but just because they're both called "gpt-4 something" doesn't mean they're the same under the hood or that they didn't change a bunch of stuff under the hood to make it cheaper to serve
Market discovery
Why do prescription medications cost so much, and generics so little (comparatively)?
Artificial inflation to recoup R&D.
Its very clear: nobody wanted to pay for usage at that price point
Why is this so difficult to understand?
1. the field was nascent and new efficiencies were discovered
2. supply and demand
3. its in the company's incentive to make their models more efficient to increase overall usage so that while the margin remains the same, the total revenue + profit increases
I genuinely don't know what puzzles everyone?
Yeah I've been thinking about this and the analogy I came up with is that tokens are basically equivalent to an in-game currency in -free to play" mobile games. You're trading actual money for some notional "curency" or coins that can only be used for one thing but, unlike mobile game coins, you don't know how many coins something costs before you use them. It's kinda weird.
Because as per usual it's silicon valley misunderstanding economics. AI is HPC. And how the HPC market worked before:
If you're the best performing "computing cluster" (ie. whatever you call the entity that can complete a massive calculation), you get a blank check from Congress.
Why? Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity. And of course, they were replicated worldwide for this reason. I mean not that anyone will admit this but we don't have the best possible solution, and we don't know either the upper or lower limits for fusion devices (plus the lower limit would be very useful for energy generation, which for the US would effectively mean almost literally unlimited large marine ships that never need refueling. And yes, the solution to that problem is almost literally a 3d shape. Not just that, but mostly)
Now it appears it does not work the same when you democratize computation. Humans want a particular amount of computation and are willing to pay a given price for that. But the accountants still saw the blank check from before and ... do what accountants do. Economics don't change because you make things bigger and accessible, do they? Oh ... wait a second ...
As someone put it recently though, we now have data. 2.3% of humans in the US are willing to pay $20 per month for the support of a model like GPT-5.5/Claude code. If that's true (and after years of having this model, why wouldn't it be?) ... it means AI startups are doomed (because it's not even 10% of what they need it to be to make economic sense).
> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference.
The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens.
Everyone is trying different pricing schemes and discounts as they test the market. The demand is fluctuating at the same time.
It’s probably very confusing if you’re primarily familiar with stable and mature markets. Price fluctuations are a common feature of new and evolving markets.