logoalt Hacker News

latentsea • today at 4:19 PM • 0 replies • view on HN

You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.