I find that fp8 cache can be pretty bad in vllm but works fine in llama.cpp. I don't know why, ...

beacon294 • yesterday at 8:29 PM • 1 reply • view on HN

I find that fp8 cache can be pretty bad in vllm but works fine in llama.cpp. I don't know why, but I plan to review the implementations.

Replies

CMay • yesterday at 10:03 PM

Llama.cpp implemented some rotation optimizations for quantized kv cache to improve the preservation of attention quality or similar, after everyone was talking about TurboQuant. It's not perfect and when you're talking about long form reasoning, little differences can make or break the results so it is situational.

alt Hacker News

Replies