So the question becomes, how many other parts of the inference pipeline have left 1000x optimization opportunities lying on the table?
I'm sure there's been a lot more effort put into the other, more consequential, portions of inference time.
the answer is many! This would take hours to write. Full teams and research on nearly every part. So many 'unlocks' coming.
The problem with the rest of inference is that changes are not trivially correct or incorrect, as they are with the tokenization layer.