logoalt Hacker News

RussianCowtoday at 1:07 AM1 replyview on HN

Unfortunately, the lack of an input cache discount makes it prohibitively expensive for most use cases that aren't one-shot prompts.


Replies

scosmantoday at 3:49 AM

well same applies to GPT OSS 120. Qwen is just the much smarter model of the 2 public options on Cerebras.