I thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.
llama-server also supports saving the kv cache to SSD. I had no issues with cache invalidation using pi.
llama-server also supports saving the kv cache to SSD. I had no issues with cache invalidation using pi.