logoalt Hacker News

wg0 • yesterday at 5:20 PM • 1 reply • view on HN

What's that in summary?


Replies

Wheen • yesterday at 6:02 PM

Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.

Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...

➕ show 1 reply