logoalt Hacker News

petercoopertoday at 6:10 PM0 repliesview on HN

FWIW, on my Mac Studio I get ~24-27 tok/s generation between 0-16k context in - that's on the Q6_K GGUF with speculative decoding on. I have spent zero effort optimizing/improving this so far but will be trying the 4 bit MLX next (I've tended to find models drop off somewhat below 6 bit but maybe that isn't the case nowadays).