Bit of a discount if you're using caching:
> same input and output prices, with cache reads at a quarter of the cost
This should impact any long-running agent since subsequent calls can benefit from cached reads for previous transcripts.
~30% reduction in real-world task cost vs. Fable 5 in our evals at viktor.com ! Caching goes a looong way
And yet, despite this, the quota limits went down by 17%.
~30% reduction in real-world task cost vs. Fable 5 in our evals at viktor.com ! Caching goes a looong way