If you have a DGX Spark, try my Spark/SM12x specific inference engine.
I've got it (Qwen 3.8 flash next) working (sans ... MTP working on that now).
https://github.com/rdaum/eider/
~80tok/sec prefill, 12tok/sec decode, ~80GiB memory resident, the n-gram table pages from SSD.