logoalt Hacker News

eblansheytoday at 7:33 PM1 replyview on HN

Why not just run FP8 on vLLM with that much vRAM? It's plenty fast.


Replies

hadlocktoday at 8:00 PM

For high concurrency, using the blackwell's native native W4A4 MLP compute path, nvfp4 is something like a 1.2-1.5x performance increase over FP8. We're doing data enrichment (so, tasks completed successfully + tokens/second) so the performance bump shows up in the tasks/month number.

I am just now getting the benchmarks running against 3.8 27b but I expect similar results from benching 3.6 27b at the same quant.