logoalt Hacker News

arjietoday at 5:03 PM0 repliesview on HN

The cheapest card that will run this model very well is a ln unlocked CMP 170HX. But you can run it on a 3090. I run it on an old spare A6000 Ampere. I think I wouldn’t use anything lower than 60 tok/s though, which you can get with MTP etc. I just use a full vllm stack but some people see a lot of speed with ninfer (there are non 5090 ports).

The large RAM Macs are unusable for inference of dense models as of now. Token generation is too slow.