From the looks of it, this seems focused on datacenter/batch inference, and doesn't tune its kernels to the specific hardware and workload where inference is being run like Magnitude does.
Magnitude is optimized for maximum single-session performance and memory efficiency - so we should be more performant for local inference use cases.