for 3.8? I can get 40 tok/sec on it (M3 Max 64GB), but I don't use it because I can get 90 tok/sec with 3.6 MoE (MTP + 6 bit quant), and I don't notice any quality improvements for my tasks using a dense model.
Did you ask Gemini or DeepSeek to look at your oMLX server log to see what was going on? This can help a lot if it is just a misconfiguration.