We're running Kimi 2.8 on a $107k server and getting around 50k tokens per second or better on most things.
We've already saved money compared to last years token cost on Claude/Gemini
how many requests per second can the server take?
> We're running Kimi 2.8 on a $107k server
Equipped with what? Is it CPU based inference, a mix...?
Are you developing software? Is most of it used on a coding agent? (Like Claude Code or ChatGPT Codex?) If so, what coding agent do you use? If you're not developing software what do you use it for (roughly)?
What is your config?