FYI they also released FP8 quants, and those should be faster on your setup (we have the same). As l...

NitpickLawyer • yesterday at 5:04 PM • 0 replies • view on HN

FYI they also released FP8 quants, and those should be faster on your setup (we have the same). As long as you keep kv at 16bit, FP8 should be close-to-lossless compared to 16bit, but with more context available and faster inference speed.

alt Hacker News