logoalt Hacker News

Rapzid • today at 1:18 AM • 0 replies • view on HN

During peak hours requests queue and inference slows. During off-peak they can move systems over to training.

Where is the evidence they are "nerfing" the models due to request volume?

Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.