Has anyone calculated the effective intelligence of these quantized models?
I think publishing benchmarks with quantized models should become standard practice.
There's some info in the README, including:
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
https://github.com/Niko1221/Strata#which-model-should-i-pick
See this recent paper: Quantization Degradation in Large Language Models: A Signal–Noise Perspective [1].
This repo uses 2-bit quantization and removes some of the experts for its smallest fastest model. Make of that what you will.[1]: https://arxiv.org/abs/2608.08188