Most responses here are along the lines of "model capabilites move too fast to build hardware for".
I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.
Totally. But it's worth noting that this is a pretty new thing! I wouldn't have bet on that a year ago, but now I would.
"Intelligence per second" is a striking phrase! I'll be turning it around my head at a moderate IPS until I hopefully make something of it.
But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"