logoalt Hacker News

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

114 pointsby mfiguieretoday at 1:39 AM8 commentsview on HN

Comments

arikrahmantoday at 3:12 AM

I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.

mmastractoday at 6:11 AM

I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.

TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems

vivzkestreltoday at 3:56 AM

404 on the blog page? https://zartbot.github.io/blog/

show 1 reply
N_Lenstoday at 4:40 AM

[dead]

smy20011today at 3:25 AM

Removed

show 3 replies