Can you mention what inference stack you're using? I've tried MTP several times with that ...

anon373839 • yesterday at 5:05 AM • 0 replies • view on HN

Can you mention what inference stack you're using? I've tried MTP several times with that model and it always seems to significantly cut my token generation speed from ~60 tokens/sec to ~40 (M3 Max).

alt Hacker News