Can someone please explain to me how the recent llm cache breakthrough doesn't alleviate this memory shortage issue? https://intl.cloud.baidu.com/en/article/8937874
See Jevons paradox. https://en.wikipedia.org/wiki/Jevons_paradox
I have this wild theory that this has not just todo with AI but with weapons production overall.
Not everyone is using baidu?
If training and inference is hardware constrained, and you can train and server better and bigger models on the same hardware with memory optimizations, that's exactly what I would expect companies to do.
Demand outstrips supply. Efficient software is great, but it doesn’t directly resolve insufficient supply of hardware.
Why use less memory when you can just have more LLM?