logoalt Hacker News

GLM-5.3-Flash

642 pointsby Philpaxtoday at 2:08 PM299 commentsview on HN

Comments

mmastractoday at 2:20 PM

Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash

I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.

I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.

I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

show 7 replies
bertilitoday at 4:30 PM

This is going so fast! What a time to be on hackernews:

July 16th: The "Kimi K3 moment" - China has caught up to Opus!

4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!

12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!

show 2 replies
mrngldtoday at 3:15 PM

Chinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it.

https://deepswe.datacurve.ai/

That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost.

They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts.

Congrats to them!

show 6 replies
matheusmoreiratoday at 3:25 PM

You guys read Z.ai's terms of service, right?

Broad and perpetual license over inputs and outputs, and even your name and profile picture.

Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country.

Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is.

Vague prohibitions on discussing Z.ai, even my posting this comment violates it.

Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.

show 15 replies
dzongatoday at 5:22 PM

> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified.

whether it's the cost to develop models, cost of hardware, cost of serving ie inference.

preommrtoday at 6:02 PM

So the vagueposting by googlers about Ox Alpha was just... what exactly?

Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.

show 1 reply
Tepixtoday at 7:21 PM

GLM 5.3 Flash: 320B parameters with 18B activated

Qwen 3.8 Next Flash: 125B + 51B = 176B parameters with 6B activated

DeepSeek V4 Flash: 284B with 13B activated

The new Qwen model is the most promising for one or two Strix Halo 128GB with the low number of active parameters. On paper it's much stronger than Qwen 3.8 27B.

XCSmetoday at 5:41 PM

Nice, finally they fixed the huge reasoning tokens count.

Now it's similar cost to DeepSeek v4 flash, but smarter.

My tests: https://aibenchy.com/compare/z-ai-glm-5-3-flash-max/deepseek...

sunbumtoday at 2:19 PM

> with all of this traffic served on Chinese AI chips

RIP Nivida shareholders

show 10 replies
revolvingthrowtoday at 2:21 PM

> 320B total parameters and just 18B active parameters

This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home.

@edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb.

… you’ll still need to splurge, though.

show 3 replies
pietztoday at 4:03 PM

With tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different?

Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.

show 1 reply
TaLiTrtoday at 2:24 PM

> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

From a biased source, but would be big if true. I've had great results with GLM 5.2.

From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.

show 2 replies
OldGreenYodaGPTtoday at 6:57 PM

Tested this last week and couldn't get it to finish any task that took more then an hour with /goal keep getting errors

lxetoday at 3:29 PM

Is the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is.

What irks me about this is that the harnesses seem to be just an afterthought here.

Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, etc... but when it comes to literally just setting up a productivity environment and trusting my entire machine with it, I just run Codex.

I have heard from anecdotes where people have indeed replaced their main drivers with DeepSek V4 Flash or GLM and state that "it's almost as good as... [claude/gpt]" but I never hear anyone say "yeah, this is the model/harness that I now run on my machine and don't mess with it"

show 4 replies
packetlosttoday at 2:20 PM

For those who didn't read, this is the identity of the mysterious "Ox Alpha" model

show 2 replies
cootsnucktoday at 2:49 PM

If we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute).

I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.

show 2 replies
claudeIsDowntoday at 3:24 PM

On OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M

How is the business model of Anthropic/OpenAI will sustain?

show 3 replies
yipinwongtoday at 2:46 PM

When reading this type of announcements, always have keen eyes on graphs.

e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20.

- This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports)

I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)

show 1 reply
singularity2001today at 4:02 PM

At the current 50%-off GLM-5.3-Flash price ($0.075/M input, $0.25/M output; cached input $0.015/M), surprisingly, roughly $400–900/month would buy token throughput comparable to fully exhausting Claude Max 20×

syntaxingtoday at 4:52 PM

Ironically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.

show 1 reply
iamsyrtoday at 2:18 PM

Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

- Input: $0.15 - Output: $0.50 - Cached input: $0.03

show 1 reply
mariopttoday at 2:26 PM

It's only 320B, local frontier AI is getting closer, sooner than expected.

show 2 replies
yousif_123123today at 4:44 PM

Will we need all the data centers being built or will improvements in software and hardware allow the majority of AI workloads to run locally or in the cloud but way more efficiently than was projected when all the plans were laid out?

Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?

jatinstoday at 4:50 PM

I was quite surprised that Zai had deep pockets to serve this free for a week. My first guess was this was an American lab like xai or google

BeetleBtoday at 4:18 PM

The key difference between this and all other GLM models is it's multimodal. You cannot send images to the other GLM models.

show 1 reply
garo-protoday at 2:32 PM

> Combined with our latest 30T-token multimodal pre-training corpus [...]

Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?

epolanskitoday at 2:18 PM

I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field.

It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

show 4 replies
rahimnathwanitoday at 2:17 PM

Related: https://news.ycombinator.com/item?id=49446422

(281 points, 118 comments)

AnodicElegytoday at 3:20 PM

Artificial Analysis benchmark is out: https://news.ycombinator.com/item?id=49450353

hxiitoday at 4:27 PM

In my brief testing, it did about as well as Qwen3.8-4B-Distill, and LFM2.5-2.6B overtook both.

kburmantoday at 3:35 PM

offtopic: Is there any chance we could see competing models from other countries in the next 5 years?

show 1 reply
swingboytoday at 2:32 PM

How much is the “discounted” pricing they mention?

show 1 reply
beannttoday at 5:34 PM

Is it good compare to Opus 5 ?

Destinertoday at 2:21 PM

from the article, pareto frontier for open source models is completely dominated by GLM now.

show 3 replies
jdw64today at 4:10 PM

This was the ox-alpha model, right? I remember it performed really well for a model that had 'flash' in its name.

scottfitstoday at 3:57 PM

so is it confirmed if this is the mysterious OxAlpha model?

show 1 reply
tokaitoday at 3:13 PM

Why is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.

show 1 reply
Imustaskforhelptoday at 2:34 PM

> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself.

> (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.

VirusNewbietoday at 4:51 PM

It looks like gemini 3.7 flash actually beats it in a lot of benchmarks, no?

https://x.com/Zai_org/status/2092616204787626030/photo/1

kayleykiwitoday at 2:43 PM

This looks like it goes hard, can't wait to try it

toppytoday at 2:35 PM

By clicking this link you download some PDF in the background

show 1 reply
knowaveragejoetoday at 3:57 PM

Any providers hosting it outside of China?

show 2 replies
tinyhousetoday at 3:46 PM

Anthropic is accelerating their IPO cause they know what's coming in the next 5 years.

dakollitoday at 3:34 PM

I didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.

show 2 replies
ammmwtoday at 3:08 PM

[dead]

🔗 View 1 more comment