I’m sorry but this looks like vibe coded marketing slop that’s highly inaccurate.
For one, there is zero consideration of prompt/KV caching, which we all now is basically essential especially for workloads at scale.
Secondly, it seems to base all benchmarks off a batch size of 1.
Nobody running a cluster of B200s or H100s is doing inference with batch sizes of 1.
And even worse, the calculator assumes you run models with a context window of 0 tokens? That affects how many GPUs massively.
I’m not nitpicking over small details or intentional simplifications here, but the estimates this is giving is horrendously inaccurate by a few multiples.
Thank you. For me the hint was: “How many GPUs is 1M/B/T tokens?“
Those units…
Also “is” is load-bearing
What! If you start running multiple models through the same GPU, you'll end up with cache issues like spectre on Intel CPUs!
This is how viruses swap DNA, and it's how AGI and sentience will accidentally happen. A mix of code here, a combine of cache data there, and AGI inadvertently appears!
I don't let anyone in my company do this, and you should not do so either.
It's all fun and games until you spark the literal apocalypse, dannyw!
So yes, the website is 100% accurate for safe LLM usage.