The speed at which tokens are crunched, even on the same hardware, differs between models as well. Using more tokens is only a problem if they are processed at the same speed as with a comparison model.
Using more tokens is a significant problem if you pay per token?
Using more tokens is a significant problem if you pay per token?