logoalt Hacker News

GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost

219 pointsby ed-is-aitoday at 4:24 PM91 commentsview on HN

Comments

jchwtoday at 4:56 PM

Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.

"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...

This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.

show 1 reply
hellohello2today at 5:24 PM

This whole thing immediately reads as Claude generated, making it hard to take seriously.

Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent

show 2 replies
gertlabstoday at 5:17 PM

One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.

We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.

Data at https://gertlabs.com/rankings

iamcoder18today at 4:56 PM

There's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.

solenoid0937today at 5:56 PM

IDK I use open models every day for personal projects, and closed models for work.

Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me.

I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.

ac29today at 4:54 PM

Not sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last place

show 2 replies
Aeroitoday at 6:25 PM

ai;dr

can't take any generated benchmark seriously. if you produce actual results, then produce actual copy to go with it.

CMaytoday at 5:07 PM

> if you run one model, run glm-5.3

That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.

Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.

Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.

nylonstrungtoday at 6:45 PM

The stealth model "Ox Alpha" has been crushing benchmarks and appears to be the next release in the GLM family

Klaster_1today at 4:59 PM

For the last week, I've been heavily immersed reverse engineering a device with help of GLM-5.3 and it surpassed all my expectations - I actually managed to achieve very way more than I thought I would. I never worked on such low level stuff, it would have taken me months to learn ARM assembly and how to find for and write exploits. Initially, I attempted this with Claude, but it blocked me on the very first message, so I got a refund and decided to try z.ai. The only downsides are that it's maybe a bit slower than my day job Opus and I had to pay ~200 EUR for a monthly plan in order not to bump into weekly limits in a couple of days. If this level of capability cost maybe 50 EUR, I'd strongly consider getting a long time subscription.

show 1 reply
pulkitsh1234today at 7:12 PM

>Short version: if you run one model, run glm-5.3 — 100% pass, a 9.3 rubric, $0.28 for the lap, about a fifth of gpt-5.5's cost

... I mean wtf is this prose ? what rubric ? what lap ? I can definitely say this is Opus 5.

scottfitstoday at 4:52 PM

my prediction is even if open source Chinese models are 90% as good (or even a bit better, which I don’t really believe because of benchmark hacking) enterprises will still pay for Claude / ChatGPT and the harness, integrations, and peace of mind versus using some Chinese cloud.

show 4 replies
gilesvangruisentoday at 5:39 PM

All of the answers. None of the understanding.

Arcurutoday at 5:38 PM

Anecdotal, but from my personal usage I found GLM-5.3 was not as capable at performing autonomous tasks as Opus/Sol. Certainly competitive with the Sonnet/Terra level, but not with the Frontier.

I got their lowest subscription tier and burned a week of quota on trialling it.

svachalektoday at 4:55 PM

Nice to see the TTFT chart, wish aggregators like OpenRouter would track this. Matches my experience, the Deepseek models while fast overall can have a horrendous wait before they start responding, and Claude models are superbly responsive. It's particularly annoying that models like flash and luna, where you've explicitly chosen speed over quality, can still stall out before they even get started.

_joeltoday at 5:01 PM

Sorry, just can't read that page, too AI spammy

zero0529today at 5:22 PM

Having used GLM-5.3 I honestly don't think it is better than 5.1. It is slower and the result is often overengineered, it is if it overthinks everything.

silverwindtoday at 5:07 PM

Those benchmarks don't tell much, they only check if a problem was solved, not how. Also there's surely a lot of benchmaxxing going on in the model training.

visiondudetoday at 5:27 PM

so the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task list.

gosolozerotoday at 5:03 PM

2 articles on how G 5.3 is the best in the top 10? Seems like a bit of astroturfing going on

rfgplktoday at 5:30 PM

In their current form open weight models are simply not worth running. Literally the amortized cost of hardware + electricity you need to operate them is >> than the cost of paying for subscriptions.

dainiussetoday at 5:29 PM

Is the beater in the room with us now?

pcweldertoday at 5:33 PM

The whole thing (article, benchmark) is a slop soup.

- Metaphor overload. We get it, it's just like car racing. Show some mercy on human readers.

- X, not Y

- A, never B

- Tasteless em dashes

- Hallucinated data, like model add date. The standings are so unbelievable that they border on laughable.

angoragoatstoday at 5:22 PM

The writing in the first three sentences is so bad that I closed the page.

Please, bloggers, write with your own voice. Don’t let an LLM do it for you.

hereme888today at 5:36 PM

Chinese propaganda. Both current top articles on HN are shilling for GLM-5.3

couAUIAtoday at 5:33 PM

haiku 4.5 top 7, that benchmark is absolute crap

tamimiotoday at 6:01 PM

What’s the best model for planning and architecture design, rather than solving problems in codes or an issue?

CamperBob2today at 4:27 PM

The GLM-5.3 weights are not yet open, and they've said that the delay is due to the need to nerf them for "safety."

So I have a feeling a lot of these early claims are not going to pan out in the long run.

show 1 reply
amazingamazingtoday at 5:43 PM

Hopefully all of these models lead to manufacturing breakthrough so we can bring down prices of cards.

sehwtoday at 4:52 PM

open-source when?

jacobgoldtoday at 5:41 PM

The entirety of the coding evals are just 7 trivial coding tasks in Python? This is a joke.

spiderfarmertoday at 5:27 PM

And now there's 0x Alpha, which is their (now free) new model.

ed-is-aitoday at 4:24 PM

[flagged]

qwertoxtoday at 6:07 PM

I can't wait for Mistral to host these models. For those who don't know yet, Mistral is pivoting to also offer Chinese models in their own cloud.