Is there a world where we look back at how AI usage is charged today and we equate it with how we had minutes on AOL and how absurd it seems looking back?
I think without a doubt that will be the case, unless the trend of compute getting better over the last 60 years suddenly stops. It should become less tough to run a local model and our devices should become more powerful.
However I think there is still a significant runway for these models to scale, so there will always be some sort of offering from providers. I can't imagine that our current use of the context window will be how that looks in a handful of years.
Not when you consider the sheer amount of compute to calculate a single token. Today, at least.
Maybe if we have some breakthrough in how to get equivalent ai capacity out of less compute it’ll seem crazy in retrospect.
But the magnitude of flops per token on SOTA models is mind boggling huge.
Absolutely. Tokens will flow like electricity, in the instances when you aren't just paying for electricity to run models locally. The token providers (well, at least the frontier labs) will fight it tooth-and-nail with the example of telecoms being "how to become a commodity" but as long as open models continue they will have no choice.
Any other outcome would be pretty terrible.
AOL minutes seem antiquated, but not absurd.
Eventually.
Absolutely. LLM providers are likely to be the next telecom providers. Low and steady income, translating to safe dividend income for investors. Most of these companies have no moat, since most LLM implementation strategies more or less converge to the same result.