logoalt Hacker News

Sub-1-Bit LLM Compression via Latent Factorization

56 points • by brainless • today at 1:29 PM • 8 comments • view on HN

Comments

augment_me • today at 4:53 PM

Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.

I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance

big-chungus4 • today at 4:28 PM

Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory

badatnames • today at 4:24 PM

Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means

bArray • today at 5:09 PM

Has anybody tested this? Are there any available computed models to test?

nbutton762 • today at 4:46 PM

Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!

nico • today at 4:26 PM

Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?