logoalt Hacker News

porphyratoday at 6:41 PM3 repliesview on HN

Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?


Replies

gardnrtoday at 6:44 PM

They make a giant inference chip. Their inference service is basically just advertising for their core value prop: hardware.

The CEO was on Gradient Dissent a couple years ago: https://www.youtube.com/watch?v=qNXebAQ6igs

codexontoday at 6:51 PM

The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).

show 2 replies
altertabletoday at 6:42 PM

Mostly economics I'm sure