logoalt Hacker News

Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

203 pointsby garo-protoday at 11:49 AM86 commentsview on HN

Comments

SwellJoetoday at 2:42 PM

Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio.

I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context.

And, MoE should make it run at a close to usable speed.

show 2 replies
ddtaylortoday at 1:43 PM

I enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful.

OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win.

However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame.

I tried to get in contact with them at OpenRouter about this and I was interested in working with them in the past, but it's difficult to get in touch with the right people and they are growing very fast. I expect being acquired by Stripe will accelerate those problems in some ways. I have no doubt they will resolve all of these issues eventually and scaling that much that quickly is really hard, so kudos to them, but the road has been pretty lame and taken some wind out of my sails.

show 4 replies
notnullorvoidtoday at 2:16 PM

It will be interesting to see the intersection of this with inference engines like FreeToken which improve distribution of work for MoE models across CPU/RAM and GPU/VRAM.

If all it takes for a competitive model to run locally at good speeds is a used 3090 and some DDR4, then we might be in for the year of local AI.

https://github.com/FlashML-org/FreeToken

show 2 replies
syntaxingtoday at 2:53 PM

Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.

show 1 reply
big-chungus4today at 1:38 PM

> We are releasing these architectural improvements ahead of time so that the community can prepare for the upcoming full family of Qwen4 models.

That gives me hope that "full family" means it will include smaller models like 4B.

c16today at 3:19 PM

+1 to the long list of people hoping for Qwen3.8-27b A3B.

show 2 replies
pwythontoday at 1:38 PM

I was already rolling around the idea of a 128GB M5 Max MBP. Now this!

A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.

show 4 replies
hedoratoday at 2:15 PM

Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days).

Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory.

I’m wondering if it can replace claude for llm-friendly coding tasks.

show 1 reply
Catloafdevtoday at 3:09 PM

Very curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse.

show 2 replies
honestlyrankedtoday at 1:26 PM

Alibaba is giving sleepless nights to the tech giants

show 1 reply
freddiehdxdtoday at 4:35 PM

Does it support vision?

fkndkfntoday at 3:55 PM

I can feel Dario Amodei's tears in the announcement :)

big-chungus4today at 1:33 PM

I hope there is going to be a free endpoint... Unlike 35B-A3B, I am nowhere close to running it locally

bellowsgulchtoday at 1:51 PM

Really happy for those with 128GB+ RAM. Sitting here with my Apple M1 Max with 64GB though. Was looking forward to a Qwen3.8-35B-A3B like many others.

show 1 reply
louskentoday at 3:19 PM

gpt oss killer? this can easily run on a server cpu with its memory bandwidth

show 2 replies
dmeadtoday at 2:55 PM

This is great. I have a weird system layout (192gb system ram, 8gb vram). the mixture of experts models have been nice when i can run the dense reasoning layers on the gpu (which somehow fit?!) and then the expert on the cpu.

its worked out to to 40 tokens/seconds on their 80b-a3b model. we'll see how much of a hit this is.

cogman10today at 1:29 PM

Wow. I wasn't expecting this. I thought they were going to do a 35B model instead.

show 1 reply
Alien1Beingtoday at 4:19 PM

AI VENDOR PRESS RELEASE

isattytoday at 2:38 PM

Can I run a fp8 quant with 96gb VRAM?

show 1 reply
BrucecarlLtoday at 1:44 PM

Waiting for the performance report! Ai hope it can beat DS

tarrudatoday at 12:00 PM

Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page.

show 2 replies
tw1984today at 1:39 PM

Qwen4 sounds exciting

cyanydeeztoday at 3:09 PM

oooh, I like a6b; that will be nice. 3.5 A10B qwen works really well in deer-flow when you want to seriously vibe code or research and you're just not going to baby sit.

onesandofgraintoday at 2:50 PM

where are the humans geez

blurbleblurbletoday at 1:34 PM

gg

show 1 reply
mrdoetoday at 1:29 PM

lol blocked with dns4eu

what a joke this resolver has become

metrofuntoday at 2:07 PM

[dead]