logoalt Hacker News

codexontoday at 6:51 PM2 repliesview on HN

The wafer only has space for 44 gb of sram. If they offload ram they lose the speedup of having everything on 1 chip (the whole point of cerebras).


Replies

porphyratoday at 6:56 PM

They can host larger models by pipelining it on multiple wafers. Each wafer stores one layer and N layers can serve an N * 44 gb model with N concurrency. The limitation would of course be inter-wafer I/O, which my comment was getting at. That's probably how they can serve bigger models like GPT 5.6 Sol [1].

[1] https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf...

show 1 reply
minimaltomtoday at 7:34 PM

[dead]