logoalt Hacker News

anana_today at 5:30 PM6 repliesview on HN

For more context, this puts it on par with models like GLM 5.2 and GPT 5.6 Luna, which are far larger


Replies

anana_today at 6:22 PM

And to read the tea leaves a little:

3.8 actually performs slightly worse than 3.6 on AA-Omniscience Accuracy, which could imply that they traded out world knowledge for capability in other areas.

It also produces nearly twice as many tokens per task as 3.6 (and by extension, time), which may be a tradeoff required to achieve correctness at this parameter size.

show 1 reply
bertilitoday at 6:17 PM

And more context:

Same score as the latest DeepSeek Flash 0731 which has 284B parameters! (13B active)

Its also the second best Qwen model, much better than Qwen 3.7 Max, but significantly below Qwen 3.8 Max.

show 1 reply
nsingh2today at 6:23 PM

Also with Qwen 3.8 being more token hungry than Luna, using around 2.3x tokens. Which hurts for local deployment.

show 2 replies
johnnyApplePRNGtoday at 6:23 PM

We don't actually know how large they are, actually.

show 1 reply
catigulatoday at 7:35 PM

Which should tell you how useful these benchmarks are.