The token price isn't the only reason to run a model locally though. You can do additional training to specialize or remove censorship that may be a no-no per TOS with cloud GPUs.
Cloud GPUs have ToS about purposes you’re allowed to crank numbers for?
I thought this only applies to LLM inference providers, but not raw GPU rentals.
Cloud GPUs have ToS about purposes you’re allowed to crank numbers for?
I thought this only applies to LLM inference providers, but not raw GPU rentals.