logoalt Hacker News

kamranjontoday at 1:36 PM6 repliesview on HN

Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperforms GLM 5.2 on nearly every metric. Thanks for sharing the news, I was refreshing huggingface but gave up thinking it likely would take some more time.


Replies

toughtoday at 4:06 PM

I've put Opus to it and it says it will take 3-4h to do the process to the new weights. hoping it works!

vmt-mantoday at 2:07 PM

are you working in earplugs? :)) even with 128 gigs of ram it must be super noisy.

show 4 replies
ycui7today at 4:07 PM

you don’t need new weight. try vllm-moet from github. it will autogenerate 2-bit plane.

theturtletalkstoday at 1:58 PM

What kind of tps are you getting?

show 1 reply
knuppartoday at 3:54 PM

same. been running it in a dgx spark and it slaps