Exactly, imo it’s not even that cheap if you look into perspective and consider the fact that providers could subsidize the cost of cached input tokens to virtually zero if they would allow for a more flexible API (e.g. tree of message blocks instead of chain). Most of the cost is the infrastructure around keeping KV caches, estimating their lifetimes, etc. When mist people just want to run one context block with multiple subsequent variants of a second block in parallel. I still stand by my statement.
That's an interesting point, if you send a batch with a shared prefix you basically only end up paying for the sequence length difference effectively.
There is still some minor memory bandwidth issue on outputting more tokens, but the truth is that if you process e.g. 16 messages at once you wont end up being much slower than Jev even though you have to perform several autoregressive passes.