logoalt Hacker News

girvo • yesterday at 8:52 PM • 1 reply • view on HN

That’s fascinating, but not that surprising to me. We act like quantisation is free “Q8 is basically lossless” is often said in the local LLM community, but it really isn’t. The trade offs are worth it, personally, and the damage to coding ability seems low: decision model approaches are stricter though

Super cool finding!


Replies

ByteAtATime • yesterday at 9:39 PM

Interesting - I wonder if it's because coding doesn't use the specific token probabilities, while decision models do

➕ show 2 replies