One could also run it locally on a used dual xeon (or amd-equivalent) server with 512GB RAM, albeit slower, if you have a useful workflow for it that's like "take this day's efforts and run it through various analysis agents", combined with giving it one-shot tasks/modules to build overnight. You would want a place like a garage or basement to put the server because it'll be loud.
> "dual xeon"
Does inference make full use of the memory bandwidth in a NUMA system?
You'd also likely spend far more in electricity than the API cost of processing the prompt(s)