logoalt Hacker News

prometheus1992today at 5:40 PM2 repliesview on HN

It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat


Replies

Baloogatoday at 7:05 PM

Now I'm feeling pretty good about getting 10-11 tokens/sec running Qwopus 3.6-35B-A3B Q6_K on an old Mac Pro 2013 (trashcan) with 128GB RAM (DDR3), 12 core Xeon, dual D700s. Arch Linux and llama.cpp.

show 1 reply
carloslfutoday at 6:38 PM

interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out!