logoalt Hacker News

skeledrewyesterday at 6:59 PM2 repliesview on HN

> $20 gets you nothing beyond what a local model running on your own graphics card can deliver.

I'd guess you're deliberately exaggerating here, but still. I've never clocked the actual tokens/second, but I'm on the $20 plan and get ~15M tokens/month for fully utilized weekly quotas (checked couple months ago). Meanwhile the best I've been able to get locally was ~8 tokens/second with Qwen3.6 35B A3B, which is wildly painful for coding sessions and gets a maximum ~20M tokens in a month... if it's going 24/7.

Just wanted to stick some empirical data here, given that statement.


Replies

bot403yesterday at 7:28 PM

I run local models. Your op is absolutely wrong. To get a local LLM is at least a $1500 investment at the cheapest. $5000 if you want usable.

At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.

show 3 replies
nozzlegearyesterday at 8:38 PM

Ouch. I get ~45t/s with Qwen3.6 35B A3B, and around ~70-80 with my current model Ornith 1.5 35B A3B. Local models work a treat IMO if you've got decent hardware for it.