logoalt Hacker News

ryeguytoday at 3:57 AM1 replyview on HN

I keep seeing mention of the cache, what's special about it? All frontier llms have prefix caching, what is special about deepseek's approach?


Replies

fspeechtoday at 7:20 AM

Their kv cache is smaller so they can use less vram and also keep your prefix cached for longer. https://deepseek.ai/blog/deepseek-v4-compressed-attention

show 1 reply