Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
Yes, but… more thinking tokens also means longer solution generation time. That said, v4 Flash is a fast model. I use it all the time because it’s very smart for the price. But it is verbose sometimes.
It’s not outdated at all to use tokens to estimate performance, it’s directly related.