I've been having similar thoughts, regarding the gigantic trillion parameter models. I'm starting to believe the future will be very specialized focused models thant can be run on modest hardware (locally) but that can scale in performance (latency, speed) in the cloud, much like any other software of today.
If you need to do programming do we really need trillions sized models? Other domains might be large or smaller, but there's no need for a model to 'know' everything and datacenter levels of hardware to run.
General chatbots might work better as larger models since you really don't know what people will also for, or alternatively we find a way to route the initial question to the appropriate model. Like MoE but without needing to load a gigantic model into memory first.