Probably. I've spent zero time with optimization at this point. Code is all new this morning.
Curious if llama.cpp is doing full BF16 for the n-gram embeddings table, or if that's quantized, too. I was going to try nvfp4 for that but wasn't sure about the quality risk.
Spark usually wins on prefill, not decode. I think memory bandwidth is about same between the two. What do you get for prefill/prompt?
Ok, at 50k context its about 126 prefill, 13 generation.