logoalt Hacker News

throwitaway222yesterday at 5:11 PM4 repliesview on HN

If companies are really doing this, then we're saying they have no problems spending tens of millions to get somewhat decent TPS and then having their employees complain they are timesliced and getting lots of timeouts because their org has 500 employees?


Replies

vohkyesterday at 5:30 PM

This is all spitballing, but I'd wager it's a third the type of workloads, a third hedging against your business depending on a single external provider, and a third trust.

Not everybody is coding or doing work that lends itself to burning tokens for warmth. Reuters for example seems to be more interested in using it for research, editing and formatting citations and the like. There's only so much of that work that needs doing, it doesn't always need to be real-time, and they probably don't see it scaling exponentially. They also need to be very aware and in control of their model's biases, or they risk it compromising their work output.

It's widely expected that all of the major providers will need to - and surely want to - drastically raise prices to justify the ludicrous amount of capital they're burning. Multiple companies have already talked about how their AI costs have exploded, and from what I understand that scale of enterprise is paying API rates. I would be disappointed if big business wasn't having a think about what that liability could look like. It's one thing to be reliant on a relatively "stable" vendor like Microsoft for Windows and Office, another to get AWS sticker shock, and then this is promising to be an order of magnitude worse.

Then just plain trust. What if ChatGPT starts recommending your competitors products, or the USA bars export of Anthropic's latest model (again, but for real this time), or they stop serving a model your business now depends on, and so on... That's a lot of risk to leave outside of your control.

usrnmyesterday at 5:15 PM

Open weights != self hosted

WarmWashyesterday at 6:44 PM

At least accord to the MIT study last year, most employees are using their own AI subscriptions to do work.

Keep in mind that a vanishingly small number of workers are SWE's churning millions of tokens daily.

show 1 reply
kittikittiyesterday at 6:01 PM

I'm not sure why someone hasn't developed a company offering services that distributes AI across all idle or under-utilized VM's and PC's for enterprises in order to serve open sourced models. Outside of the electricity bill, there's no additional expenditure and you get the AI.

We've all seen the office spaces where there's 200 empty computers on a floor. Combined, it's something like 500 cores at ~3 Ghz each and around 3 TB of RAM. The networking is already there and software like exo already exists.

show 2 replies