logoalt Hacker News

verdvermyesterday at 10:11 PM0 repliesview on HN

Quantization is typically very cheap and fast. It can even be done on hardware that does not fit the model, by processing the weights layer by layer.

I use this project: https://github.com/vllm-project/llm-compressor