That's an interesting point, if you send a batch with a shared prefix you basically only end up paying for the sequence length difference effectively.
There is still some minor memory bandwidth issue on outputting more tokens, but the truth is that if you process e.g. 16 messages at once you wont end up being much slower than Jev even though you have to perform several autoregressive passes.