logoalt Hacker News

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

327 pointsby altertabletoday at 6:32 PM106 commentsview on HN

Comments

nostreboredtoday at 7:29 PM

150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool.

Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information.

``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```

We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives:

``` {"message":"Model does not exist or you do not have access to it.","type":"not_found_error","param":"model","code":"model_not_found"} ```

When the error is really about billing.

I always want to like Cerebras, but I get the vibe that as a tokens in tokens out consumer you are not valued at all.

show 1 reply
jasongilltoday at 6:50 PM

It would be great if they made their inference capacity for this model available via OpenRouter; the fastest provider on OpenRouter right now is at ~80tps https://openrouter.ai/qwen/qwen3.8-27b#providers

They do appear to host other models on OpenRouter so maybe Qwen3.8 will be there soon: https://openrouter.ai/provider/cerebras

show 1 reply
gpugregtoday at 7:34 PM

I was wondering whether this was any good for programming, but it is too fast for its own good. There is a limit of 450,000 tokens per minute. I hit this limit in about 90 seconds and burned through $1.10 while doing so. This is because cached tokens count towards the token limit.

For comparison, I ran the same task with DeepSeek-V4-Flash, which finished in 172 seconds and cost $0.024 with a final context window size of 55217 tokens, while Qwen3.8-27B was not even close to being done with a 64178 context window.

This is a very efficient way to burn your money, but I would not recommend it for programming.

On the positive side, I got a $5 signup bonus, so it wasn't my own money.

hexa00today at 7:23 PM

Just tried it on a medium size coding/debug problem on an existing codebase, observations: - Input doesn't look faster than other models, it spends a lot of time reading Read about 5M tokens - Output is awesome, super fast as you expect from the 1500t/sec I think that's correct - Tool call is failing more than say DS4, which leads to time wasted on retries (complex tools like browser control for example) - Shell commands are still somewhat of a bottleneck

The net effect is that I spend about the same time waiting, and I still need to read that output so, at least for coding, it actually reconciles me with the 100-200t/sec you can get on DS4 or the like. Maybe that's a good sweet spot after all and faster t/sec is not where the bottleneck is.

Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy

show 1 reply
pllbnktoday at 7:21 PM

Just a couple days ago I learned about ninfer (https://github.com/Neroued/ninfer) and on RTX 5090 I can now get ~200 tok/s and over 400 tok/s on concurrent requests which is plenty fast for a local model of this strength.

show 1 reply
gardnrtoday at 6:37 PM

I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far.

Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.

show 5 replies
taconetoday at 6:48 PM

Noticed they are present in OpenRouter, but Qwen 3.8 is not there yet. Hopefully it'll get there soon.

For those who haven't noticed though, the context size they allow for Qwen is just 128k. Still interesting as a specialized sub-agent but not really well suited for long tasks.

show 1 reply
RomanPushkintoday at 9:06 PM

The question is whether Cerebras is available... I've been trying to get https://www.cerebras.ai/code for at least 1 year now. It's all sold out. Always. I once joined their Discord, waited for the drop, and it all sold out in seconds. I haven't had enough time to put my card details. Somebody recommended that I should put my card details in advance, lol.

The next time I hear about them I am laughing, because when I could enjoy these powers? How many years I should be sitting in a waitlist...

elitoday at 7:48 PM

I just did a little anecdotal test. Had pi + cerebras review a recent commit and asked a few quick followups on it. Worked great.

The Cerebras session cost me $1.60 and took a total of 5.1 mins. I did get a few brief 429 rate limit errors in there. The p50 speed was 890 tok/s and 0.64s TTFT.

Using OpenRouter averages, that would've cost $0.29 (no cache discount at Cerebras!) and would've taken about 14.4 minutes.

So on this one short session, cerebras was 5.6x more expensive in exchange for being 2.8x faster. Or, another way, $1.32 buys back about 9 minutes of your time. Not a bad trade IMHO but the cache situation is a real bummer. The longer your session the more relatively expensive Cerebras gets. The "good" news is you're also limited by its short context window.

(Also, I used to be on the Cerebras coding plan and the support is pretty bad for end users. My guess is these public endpoints are really just product demos for potential enterprise customers.)

show 1 reply
freehorsetoday at 7:06 PM

I have used their gemma 4 31b model through kagi and getting real instantaneous answers is absolutely crazy. A very different feeling and UX. Even if the model is smaller, there is definitely a use case for these. I was wondering if they would put the qwen 27b model, it sounds very interesting to try.

show 1 reply
codazodatoday at 9:07 PM

Do I understand their pricing correctly? This is $10 per month for a developer account PLUS you pay $1.49/M for output tokens and $0.99/M for input tokens on Qwen 3.8 27b with a 128k context?

EDIT: Or, maybe it's just token pricing, but $10 is the minimum? Maybe it's that.

https://www.cerebras.ai/pricing

show 1 reply
foundfontictoday at 6:46 PM

I really wish they had their customer support somewhere else than Discord, which seems to think I'm a bot and doesen't accept my email or phone numbe

show 1 reply
orliesaurustoday at 7:55 PM

Qwen 3.8 27B is an exceptional model for coding and ranks as one of the best local models for coding....BUT in my head I am confused why a company that's IPO'd doesn't invest in RL'd super specialized, super-damn-fast models for very specific tasks - instead of giving us the OSS GPT model from what feels like 200 years ago

show 2 replies
ecshafertoday at 8:05 PM

I have a self hosted Qwen 3.8 27B and I find it to be unusably bad. Using it agentically, it will spin around in circles on even small tasks talking to itself until it loses context and starts again. I even had it say "I've forgotten the users initial question"

show 2 replies
byakotoday at 6:54 PM

1500 tok/s is wild. meanwhile my brain does like 2 tokens per minute and half of them are 'uh'. we have truly reached the singularity

show 3 replies
dshattoday at 7:02 PM

I'm saddened that Gemma4 is replaced by Qwen 3.8 on PayGo plan. Gemma4 31B is not coding model but it is excellent at intent understanding and task execution used in agentic software. This just shows that real world dominant usage for llms so far is to code generate. And not to augment business products. They must had barely anyone using Gemma to remove it from that tier.

peri-cltoday at 6:47 PM

(Was anyone able to create an account just now? I tried but onboarding falls into a redirect loop)

(update: I got my answer. support@ replied and said my email domain is on their blacklist. It was just me (and I've resolved it)).

show 1 reply
gravtoday at 7:55 PM

Should be available in OpenCode once this lands: https://github.com/anomalyco/models.dev/pull/6199/changes

show 1 reply
porphyratoday at 6:41 PM

Why do they only host small models rather than the 2.4T version? Is the I/O and interconnect between the wafers bad due to the limited beachfront relative to the massive size of the chip?

show 3 replies
the_duketoday at 7:09 PM

Funnily enough the pricing isn't that much worse than on openrouter, where the best price at the moment is $0.24 in / $2.55 out, vs $1 / $1.5 on Cerebras.

Sure, 4x input , but cheaper output. Though Cerebras doesn't have prompt caching, so not great for agentic workloads. (they do, but it doesn't affect the price.

show 1 reply
darkbatmantoday at 7:09 PM

I have been their user for more than year even used coding plans, though for normal coding the quota will definitely be a blocker if you are using opencode because rpm are bit less. Good for products/api though.

polygottoday at 6:59 PM

Ut oh, might be down: "Unable to connect to the server. Please check your connection and try again." when sending a message to Qwen 3.8 27B.

vb-8448today at 7:00 PM

At that speed it's too pricey for agentinc tasks.

show 1 reply
fulafeltoday at 7:06 PM

What are the best benchmarks/leaderboards that compare task completion time between provider+model combos?

srcreightoday at 7:39 PM

How many years until chips like this are available to consumers?

show 1 reply
drchaimtoday at 7:12 PM

The idea of custom software on the fly is coming

Marciplantoday at 6:43 PM

used their Code product with GLM4.7. its fun but if the model is bad it just doesn’t do much useful.

Hope they add such models to Code too :)

show 1 reply
trvztoday at 6:48 PM

Normal people: tok/s or t/s

Psychopaths: tok/SEC

show 2 replies