logoalt Hacker News

Infernaltoday at 4:20 PM1 replyview on HN

I am really surprised you’re seeing half the generation performance of 3.6 with 3.8 at the same parameter size and quantization (and same prefill performance to boot) - is there just an optimization in the stack somewhere for 3.6 that hasn’t landed yet for 3.8?

Curious if other folks are also surprised by this apparent discrepancy or have a ready explanation.


Replies

woadwarrior01today at 4:25 PM

I suspect this might be due to MTP mispedictions. Either because the 3.8 quantized model weights do not include MTP heads, or as what happened with Ornith-1.5 recently, corrupted MTP heads, or due to a software issue in Ollama.

show 1 reply