Output from Cerebras with GPT model is 750 tokens per second.
Don’t blink.
(Chatjimmy has 14,200 TPS.)
ChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck.
Unfortunately AMD bought them, so I don't think we will get to see another release from them.
ChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck.