I totally agree with what you're saying. Your point about other methods to leverage computation is true and important.
But that's exactly my claim: Smaller models can win on the computational efficiency front, but not the overall capability front. And as long as compute is getting cheaper, investing in more efficiency while there are still major capability on the table isn't a good business strategy.
Smaller, efficient models could lead to some really interesting things though especially considering it could lead to some Jevon's paradox like moment. To be honest, I feel like the biggest issue in regards to LLM usage in practice is that the patterns of use aren't really well developed. Yes we have agents, but it's still somewhat unclear what an agent can "do" - People seem to mostly focus on replacing some kind of existing process with an agent driven one, but actually coming up with AI-native processes is way harder.