logoalt Hacker News

coder543today at 12:47 PM3 repliesview on HN

The weights were just released a few minutes ago: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731


Replies

kamranjontoday at 1:36 PM

Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.

show 6 replies
WithinReasontoday at 2:22 PM

Hmm, it's targeting a HW accelerator with a 128x128 matmul primitive, which one is that? Warp Group Matrix Multiply Accumulate on H100?

show 5 replies
walrus01today at 3:30 PM

the unsloth GGUF at <165GB will run on most 256GB RAM pure-CPU systems (or with llama-server and a mix of loading as much as you can onto a single 32GB, 48GB or 96GB GPU and the rest onto system DRAM).

https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF