logoalt Hacker News

pipsterwotoday at 7:03 PM1 replyview on HN

1/1000 of inference compute is a non-trivial workload at scale. Gartner estimates ~$28B in inference spend for 2026 making this a $28 million dollar per year workload (edit: based on the assumption above)

Source: https://www.gartner.com/en/newsroom/press-releases/2026-07-2...


Replies

boroboro4today at 7:30 PM

The issue is it’s cpu compute which is underutilized in gpu clusters anyway, so practically it’s not really 1/1000.

show 1 reply