logoalt Hacker News

skarztoday at 1:08 PM7 repliesview on HN

do we really need breaking news about qwen posted every single day?


Replies

KronisLVtoday at 1:22 PM

If there’s news, then yes. This is a pretty great new release for those still stuck on Qwen3.6 35B A3B if they have enough memory but don’t have super powerful compute.

I wonder if I could get this running through vLLM on 6x Nvidia L4 - the 3.6 worked great on 4 cards but sadly TP6 just isn’t a thing and I don’t have 8 cards available, maybe it’s gonna be okay with like TP2 and MTP. I have no idea at this time, probably need to test out what even might be possible.

NitpickLawyertoday at 1:26 PM

This particular release is interesting because it's a preview of qwen4 architecture. And, while benchmarks are iffy, this is a direct comparison, by the same team, with qwen3.8-27b that was pretty well received for a local model.

This "next" release adds a new concept, first public release with n-grams, I think. And it's in a MoE size that is likely to be very fast and cheap to serve (faster than 27b for sure). It's also well suited for inference on alternative compute (i.e. sparks, macs, etc) so it's relevant to local users.

pseudonytoday at 1:19 PM

I and presumably quite a few others with AMD AI or Apple Mac platforms are very impacted by this.

:)

It is very relevant and for a certain group of us, far more impactful to our work the next month(s) than any blog post could be.

toshtoday at 1:13 PM

this is a new architecture (foreshadowing qwen 4)

> trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board

https://x.com/Alibaba_Qwen/status/2092591393424515114

c16today at 1:30 PM

There are many topics, personalities and politicians we hear about daily who have no merit.

Qwen's advances do (currently) have merit.

dofmtoday at 1:29 PM

This actually is meaningful news, I think. Pretty wide audience appeal in the local LLM space too.

iAMkenoughtoday at 1:10 PM

yes there’s no shortage of online real estate