logoalt Hacker News

sergiotapiatoday at 12:13 AM1 replyview on HN

Cerebras does not share the quantization of the models so you don't know if you're getting real K3 or k3 lite or something else.


Replies

codexontoday at 7:17 AM

It most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.