logoalt Hacker News

sixtyjtoday at 8:19 PM2 repliesview on HN

Output from Cerebras with GPT model is 750 tokens per second.

Don’t blink.

(Chatjimmy has 14,200 TPS.)


Replies

tomrodtoday at 8:27 PM

ChatJimmy is a much smaller model and, AFAIK, has no reasoning capability. Absolutely insane raw speed, like a supercar, while Sol is more like a freight truck.

show 3 replies
mips_avatartoday at 9:58 PM

Unfortunately AMD bought them, so I don't think we will get to see another release from them.