logoalt Hacker News

spider-mariotoday at 3:59 PM3 repliesview on HN

> Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it.

Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB.


Replies

SwellJoetoday at 4:08 PM

Even at 30 t/s, 3.8 thinks so long, even on medium, it still takes 3x or more longer than any cloud model, in my testing.

show 1 reply
bellowsgulchtoday at 5:10 PM

Thanks, man! I’ll go use that now that I know. llama-server the last time I used it for inference with this model wasn’t able to produce work fast enough to reach those numbers.

Xeoncrosstoday at 4:01 PM

Nice, which model quantization is this? Is it on huggingface?

show 1 reply