Wait for the MTP variants that will likely be out within days. I'm on a 128GB Strix Halo box and for 3.6-27B 8bits I was getting about 9tok/sec (not great). With MTP that gets closer to 18 tok/sec (kind'a usable).
Seems like MTP is available immediately!
The MTP is available, but I'm definitely not seeing 18 t/s on the Strix Halo from the 8-bit quantization, even with MTP (more like ~10 with full context on long tasks). This is a slow model (but so was 3.6). What's your exact llama-server command that gets 18 t/s?