Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.
[0] https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise
yeah. K2.6 can run on insane speeds. So sad that they don't have K3 yet.
But it can apparently also run 5.6 Sol
7-8 figures annual spend will buy a hell of a lot of capable local inference hardware you can own, though it won't be at the absurd token/s rate, you'll be able to run almost anything on it... And it'll still have a good residual resale value after 4 years the way things are going now.