how much of performance comes from inference time tricks like scaling, topn ect . maybe models providers are also in position to run their models vs running os models by a generic providerc