Llama.cpp implemented some rotation optimizations for quantized kv cache to improve the preservation...

CMay • yesterday at 10:03 PM • 1 reply • view on HN

Llama.cpp implemented some rotation optimizations for quantized kv cache to improve the preservation of attention quality or similar, after everyone was talking about TurboQuant. It's not perfect and when you're talking about long form reasoning, little differences can make or break the results so it is situational.

Replies

beacon294 • today at 5:55 AM

I'll read it. It could be the quants too. Some quants I try are inexplicably bad, some seem better than official (or unsloth) quants... even what should be run of the mill gguf quantization.

alt Hacker News

Replies