Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.
are you working in earplugs? :)) even with 128 gigs of ram it must be super noisy.
you don’t need new weight. try vllm-moet from github. it will autogenerate 2-bit plane.
same. been running it in a dgx spark and it slaps
I've put Opus to it and it says it will take 3-4h to do the process to the new weights. hoping it works!