logoalt Hacker News

bearjawsyesterday at 11:22 PM4 repliesview on HN

If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

I've used it on a few for fun projects and its decent but the speed is crazy to watch.


Replies

conceptiontoday at 3:31 AM

It was better when they had gemma at 1k. Inco does DS flash at about 600. A few places will do K3 and GLM in the hundreds.

Such a tiny model at that t/s is less impressive than it would have been four months ago.

fordtoday at 12:44 AM

Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.

[0] https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise

show 2 replies
LoganDarktoday at 1:34 AM

Please do not try to use gpt-oss-120b over Cerebras. It is broken, screws up tool calls most of the time, forgets to end thinking blocks and has all sorts of other issues. The speed is amazing but it is absolutely not worth it, especially at that quite incredible cost. Think: $5–10/minute levels of cost with a single agent, because Cerebras also offers no cache pricing for input tokens at all.

show 2 replies
scosmantoday at 12:51 AM

Or better: Qwen 2.8 27b

show 1 reply