logoalt Hacker News

frigidwalnuttoday at 6:14 PM1 replyview on HN

Cool! I'm thinking about a local set up. What's your usual tokens/second rate?


Replies

victordstoday at 8:09 PM

Not OP, but I’m running local models on a M1 Max as well with 64GB RAM.

It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.

I’ve also used Qwen 3.8 27B but I get 10t/s on it.

It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.