Humans have human attention spans and there's a lot of study going back decades for UX design because of this. < 100 Ms is instantaneous, 100-300 ms is noticable but still responsive. At 1 second, flow.gets interrupted, 2-5 seconds, you're clearly waiting, 5-10 attention wanders and 10+ seconds, you've lost them. The 0.1 / 1 / 10 second rule comes from Jakob Nielsen's HCI work. Perceived latency matters almost as much as actual latency, which is why chat interfaces drip out/stream words instead of just dumping out the answer at the end. At 750/tok/s, for Sol grade inference, it can spend 3 seconds on thinking tokens before outputting something to the user for a better answer while still feeling usable.