This is the version we'll be testing on our rtx 6000 today! Thank you
Why not just run FP8 on vLLM with that much vRAM? It's plenty fast.
Why not just run FP8 on vLLM with that much vRAM? It's plenty fast.