Cerebras does not share the quantization of the models so you don't know if you're getting real K3 or k3 lite or something else.
It most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.
It most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.