You could run the inference locally using the open weight models. Then you'd still get to use the models without sending money overseas