How do caches work across models? I would have thought that was very model specific - if not I’ve really misunderstood what’s getting cached.
I am not sure the author of the comment you are replying to understands that LLM systems have prompt caches
I am not sure the author of the comment you are replying to understands that LLM systems have prompt caches