logoalt Hacker News

oceanplexiantoday at 4:56 PM2 repliesview on HN

Ollama? Running a q1 quant? I don't think the writer knows what the are doing here to be honest.

You will get better information cruising r/localllama for about 10 minutes.


Replies

__mharrison__today at 5:07 PM

This is a problem that Hugging Face could easily help solve.

Crawling through Reddit or forums to find the right incantation to run a model is frustrating.

bilbo0stoday at 5:02 PM

There are many, many people who are betraying a lack of facility with these models.

Why is anyone trying to run these on an M3?

I thought it was common knowledge that, if you insist on using Apple, only M5 processors and higher have matrix multiplication units in the cores?

I've seen the same thing with people buying NVidia cards with 16GB of ram and wondering why they aren't getting 100tok/s?

Guys, please, be reasonable. You'll have to get the hardware if you want to run these things fast. If you want to experiment, the slow stuff is fine. But try not to get on HN and ask why your M3 can't get 100tok/s. You're kind of out-ing yourself.