Prediction - we are going to figure out SOTA AI performance without requiring 1TB of memory within a year or so.
Of course CXMT, Micron, and family will still be profitable, but maybe not 'surge 470% from IPO' profitable.
And by the time the models that require 1TB memory will be better.
It's like people saying mobile chips are going to be better than PC chips... until they realize PC can be made of mobile chips too if that comes true.
The fundamental issue is that for these systems bigger is essentially always better (if affordable). So if we can squash something like Kimi K3 down to run on a 'normal' system, that just incentivizes devs to increase the model size until once again we're at the limit of what can be run.
What do you base this prediction on? It seems very unlikely, unless by SOTA you mean the current SOTA.
SoTA AI will move well past 1TB+ memory requirements in 1 year or so.
Also, you will be able to play with much more competent models locally. They will still feel like children compared to the adults living in the SoTA region.
Are you thinking of fundamental architectural changes (like the Transformer)? Or incremental (like MoE or GQA)? Are there specific neolabs or techs you are following that lead you to this prediction?
Indeed it seems quite possible we are one architecture breakthrough away from existing chip stock driving us all the way to ASI.
Nobody has done AFAIK the information theory to prove it’s not possible.
Yeah, you can directly print the SOTA AI model on the chip. I also believe this can be done on older process nodes. The most important factor would be speed. How fast can you go from new model to new chip?
"we" being China?
Even if we do, we are going to need a lot of RAM for the billions of agents running everywhere.
that would be perfect timing
Hyperscalers are screwed, data center mania couldn’t even be completed during this massive spending spree, all while the people seemingly sitting on the sidelines are working on getting these things baked into the OS and chipsets in a way consumers wouldn’t notice
NOPE. The opposite will happen. CXMT will flood the market thus making 1 TB models affordable.
How would we do that? There's no historic precedent for that. Its fundamentally an information theory thing: what's the max amount of intelligence you can get out of 1 KB/MB/GB? There has to be a limit and I'm not convinced that it's far off.