The model that is most optimized around token cost is, in fact, Chinese. DeepSeek is astoundingly cheap by default, but if you use it from Reasonix (the harness optimized around its cache), it becomes even cheaper.
I keep seeing mention of the cache, what's special about it? All frontier llms have prefix caching, what is special about deepseek's approach?
I keep seeing mention of the cache, what's special about it? All frontier llms have prefix caching, what is special about deepseek's approach?