Model-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference.
Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.
One thought I had is that you could use FPGAs to get hardware performance but maintain the ability to dynamically update. I don't know enough about hardware to consider trying such a thing but I'm curious if that could be made practical and economical somehow.
Many years until consumers can buy them at reasonable price. Nvdia and AMD are making GPUs bad in purpose for consumers so that nobody can build a datacenter from them. It will take a long time.
I don't think asics specific to a specific model or even model family are likely to be commodity hardware anytime soon.
It's extremely expensive to build that and you'll be at least two major model generations behind before you even get your first wafers back. By the time you got your production run ready to go and packaged for market nobody's going to care.
Once we end up going something like 24 months between major advances and capabilities for these models then I can start to see asics for a model being possible.