logoalt Hacker News

2001zhaozhaotoday at 6:18 PM0 repliesview on HN

There's now a 40X discount in the cache input pricing instead of 10X.

This seems to point to them having achieved some kind of optimization in attention mechanism perhaps along the lines of DeepSeek V4, which had a similarly high discount between cache input and normal input.

In real world use, the savings should be quite noticeable. For example, you can now use the model at 800K tokens context window at the same cost efficiency as the previous model at 200K tokens context window.