I think this whole premise is ignoring the point that most decently sized companies want to and eventually will be running their own LLM workloads. Currently use cases are limited by scope and imagination, and predominantly focused on cost efficiency. Once a use case becomes a top line revenue driver budgets will become essentially only limited by ROI. There are some inference workloads that are unacceptable to send to the frontier labs for privacy reasons. Most of it will come from firms who want to productionalize their own fine-tunes. In any case, the market for inference is less than 1% of what it will be in 5-10 years. This whole notion that GPUs will be sitting idle en masse is ridiculous. People will just start running GPU databases if it becomes cost efficient.
TIL about GPU databases. I did a bit if research but only found this one: https://github.com/bakks/virginian
What are the advantages of a GPU database?
Is anyone really running GPU databases in prod at scale? Or are you assuming that the GPU compute from a crash will become so cheap that it makes this niche scale?
> most decently sized companies want to and eventually will be running their own LLM workloads
And probably some decently sized states too. Not only commercial actors are up to the job.
[dead]
but this is just an absolute fantasy, what, exactly, will this compute be doing? Our lives are already deeply entwined with technology and use barely any compute. How could our lives become 100x more dependent on compute?
You can be thrilled by this exciting technology and all the possibilities it brings without thinking it is going to require this huge capital investment in GPUs. You could radically change billions of lives with a few dozen GPUs.