logoalt Hacker News

super_mariotoday at 1:28 PM1 replyview on HN

I was in the same situation, I used maxed out 15'' M3 Max MacBook Pro docked to Studio Display closed on vertical stand behind the screen. It was fine for office work, but running local LLMs would definitely overheat it. The battery started degrading purely due to heat issues. And it was audible as well.

I decided to get Mac Studio M4 Max, also all maxed out config and the cooling is so much better that I can run local LLMs like Gemma 3/4, gpt-oss 120b all day long without any heat issues or any audible fan noise. So for my use case it was the right decision. I subsequently added 15'' M5 Max MacBook Pro all maxed out to my collection and even though it is slightly faster on LLM inference (I get 100 tokens/s with Gemma 4 27b model), you just can't run LLMs longer than a few minutes. It starts overheating and gets really loud.


Replies

seanmcdirmidtoday at 2:43 PM

Weird, I’ve run LLM batch sessions for hours on my Max M3 MBP. It doesn’t get very loud, though I’m not getting anything close to 100 tok/s on a 27b model, I use a 35b MoE model just to get 90 tok/s. The fan comes on but thermally it never overheats. I do have it in a vertical closed position, though.