logoalt Hacker News

om8yesterday at 10:42 PM0 repliesview on HN

If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS