logoalt Hacker News

rohansood15today at 1:14 PM7 repliesview on HN

Didn't expect it to beat 3.8 27B so cleanly.

Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.


Replies

user43928today at 1:25 PM

For comparison with hosted models, GPT 5.6 Luna scores 67% on DeepSWE, compared to 59% here for Qwen.

Luna is $0.20 / $1.20 vs $0.16 / $0.47 with Qwen.

show 3 replies
overgardtoday at 4:16 PM

Curious, how are you running it and what quantization are you using? I've mostly been using MTPLX; 125B sort of looks like it'd be right at the limits of my 128GB MacBook once you factor in KV cache and context window.. wondering if it's worth it compared to the 27B model which gives me a lot of headroom or even a 72B model.

Squarextoday at 1:21 PM

I don't like these comparisons. Sure it is impressive, but it does not have a world knowledge of larger models. It has most of theirs intelligence.

show 4 replies
grueztoday at 1:24 PM

>Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

How much memory does this translate to and what quantization (if any) were applied?

show 1 reply
hedoratoday at 2:23 PM

My AMD strix halo box (haven’t benchmarked yet) should also run it reasonably well. It was $1400 at launch, and is $4K now.

Your mac is < $2K in Biden-era dollars. Presumably the economy will eventually recover; maybe in one Moore’s law doubling if the midterms go outrageously well. That’ll be two doublings since the halo launched. I’d expect this model to run on a sub $1K box by then. $2K ought to get you a 512b parameter model at that point. If we have to wait out the rest of the term, the cost cliff will be even more pronounced when it hits.

show 2 replies
dist-epochtoday at 2:02 PM

It's a much bigger model, with a next-gen architecture. It's expected to be much better.

RobertasTatoday at 1:23 PM

[flagged]