Probably because they made ASICs to run inference for less.
Are those actually deployed at scale yet?
Are those actually deployed at scale yet?