logoalt Hacker News

amelius • today at 4:39 PM • 2 replies • view on HN

For training or for inference?


Replies

anvuong • today at 6:03 PM

Both, especially for training. Astra and Fable were presumably trained on cluster of 100,000k GPUs, or at least a couple of 10Ks.

3,800 GPUs is nothing in the frontier side.

ricardobeat • today at 5:21 PM

They don't publish numbers, but Anthropic has a single DC with 200k+ GPUs for inference, GPT-6 Astra is said to have trained on 100k+ GPUs.