logoalt Hacker News

peri-cltoday at 7:29 PM2 repliesview on HN

> "Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy"

I don't believe Cerebras has a cached input pricing? They don't list one on the model page:

https://inference-docs.cerebras.ai/models/qwen-3.8-27b

edit: See the sibling discussion,

https://news.ycombinator.com/item?id=49554520#49555094 ("Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate")


Replies

hexa00today at 7:42 PM

lol yeah just saw that, yeah that makes it unusable I think at least for me.

I wonder if they will do that with sol ultrafast!

olivermutytoday at 7:35 PM

They have cache, but it costs the same indeed, no idea what the point of the cache is

show 1 reply