logoalt Hacker News

gardnrtoday at 6:37 PM5 repliesview on HN

I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far.

Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.


Replies

jasongilltoday at 6:46 PM

It appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-cachi...

show 3 replies
elitoday at 7:04 PM

Strongest model that they host on the public endpoint. They do a super fast version of GPT 5.6 Sol for OpenAI and have bigger open models on dedicated endpoints.

altertabletoday at 6:40 PM

Agreed, but in our SAAS I can tell some UX will sky-rocket to next level with this

singpolyma3today at 7:04 PM

The coding plan is gone now right?

show 1 reply
cute_boitoday at 6:47 PM

i believe they used to have monthly plan, what happened to that?