It’s interesting they skipped caching. I could see wanting to ask follow up questions so having your first x tokens in cache would be interesting.
Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.
There may well be a technical reason I don’t understand.
I'd imagine it's chasing the lowest possible latency