logoalt Hacker News

a_humeantoday at 1:45 PM1 replyview on HN

Waiting for llama.cpp support to land, but this might be a big deal for Strix Halo users.

6B active params helps around the memory bandwidth constraints, but a 128GB box can probably run the Q3/Q4 quants fairly easily with a decent context size. This might actually be better for strix users than 27B, which was already very good.


Replies

hedgehogtoday at 7:22 PM

In my early testing it's way better both quality and speed on Strix Halo (posted recipe in sibling comment).