So they quantize models, only tell about it in the blog post (instead of a warning on the model page), and even in the blog post pretend there's no difference by benchmarking on small context tasks many of which are saturated. Coding agents will probably be severely negatively affected by KV quantization.
I'd say serving quantized models without saying so on the "store" page is fraud.
don't disagree, but there is a big difference between 'quantized model/weights' and quantized activations