logoalt Hacker News

Flaviustoday at 8:10 AM3 repliesview on HN

> run free LLMs locally at native speed

This reads like a hallucination. What does native speed even mean?


Replies

kyxsctoday at 8:18 AM

for example, models running at like 100-150 tokens/second (or faster!) vs 15 t/s

(fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s)

models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB).

running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!

show 1 reply
csomartoday at 10:22 AM

There should be some kind of moratorium on new accounts. HN's always had waves of newcomers, but their impact was always limited. The wave passes and people either get filtered out or adapt. That doesn't seem to be happening anymore, since bots can churn out endless gibberish.

He did answer you though. Native is x10 the non-native speed. 50/50 that's not a bot; though it could be a meat-proxy

show 1 reply
lmpdevtoday at 8:20 AM

I assume they mean same t/sec as a SOTA cloud model