I hear the argument here, but isn't it possible it has dramatically more knowledge and when you get outside the common cases many of us use it for, it'll have completely different capabilities?
I feel like most benchmarks cluster on a reasonably limited area of human knowledge
Sort of depends on how well the core reasoning works. It’s not a big effort to connect an LLM to a search provider.
You do pay for the tokens, but in theory on a smaller model each token is cheaper.