logoalt Hacker News

toasty228today at 1:54 PM3 repliesview on HN

> absolutely no way to stop it.

Unless if they take down the flat rate subs and you have to pay api prices, usage will drop a lot. Right now I freelance using gpt, if I had to pay api prices I'd probably lose 50% of the income, if not more, with the ~40% tax on top I might as well spend my time doing something else


Replies

ekiddtoday at 2:09 PM

You're unlikely to ever be priced out of tokens, at least if you'd be willing to settle for a model closer to Sonnet 4.5. That level of model certainly isn't as efficient as Fable 5, but it can crank out CRUD apps and other consulting mainstays quite well, with some supervision.

To give you an example of model in this class, the DeepSeek V4 Flash preview is a 284B A13B model, with a native quantization mixing 4 bit and 8-bit values. You can easily run it on an RTX Pro 6000 Blackwell (or 2) at a reasonable quant, especially if you offload the MoE weights to 48-64GB of system RAM. This costs US$11,800 at Microcenter right now, and it will work in any gaming box with decent cooling and a modern 1000W power supply. Over the lifetime of the card, an entire system would cost you under $4,000/year. Power is about 300W for the card (either a blower model, or a workstation model with the power cap), and another 150W or so for the rest of the server.

Or you could buy it on Open Router from dozens of different commodity vendors, starting around $0.09 per million tokens input, $0.18 per million tokens output. This is a competitive market price, so some of the providers might be losing money or reselling surplus capacity. But given the underlying hardware costs, the numbers are in the ballpark. In other words, if you're willing to settle for lower-quality tokens, you can get roughly Sonnet 4.5 for close to free, or as a modest capital expense for a successful freelancer. Halfway decent tokens are cheap, and you can generate them in-house!

show 2 replies
LogicFailsMetoday at 2:05 PM

That would just create incentive to build better consumer HW for larger open weight models. And that would make me very happy. But I think the poster in another thread who stated there's a netflix subscription phenomena going on with the fixed rate pricing was onto something. My agents run 24/7 until I hit all my limits. I suspect my results are not typical. And as of yesterday I have issues with both major CEOs yet I thank them both for subsidizing my tokenmaxxing.

ursuscamptoday at 2:00 PM

I don't necessarily buy that API prices represent "real" prices of the models.

show 5 replies