logoalt Hacker News

jampekka • today at 9:05 AM • 1 reply • view on HN

You can get Gemma 4 26B A4B at the exact same token input price of $0.042/M. GPT-5 nano is not much more expensive at $0.05.

https://openrouter.ai/google/gemma-4-26b-a4b-it


Replies

bicsi • today at 9:33 AM

Exactly, imo it’s not even that cheap if you look into perspective and consider the fact that providers could subsidize the cost of cached input tokens to virtually zero if they would allow for a more flexible API (e.g. tree of message blocks instead of chain). Most of the cost is the infrastructure around keeping KV caches, estimating their lifetimes, etc. When mist people just want to run one context block with multiple subsequent variants of a second block in parallel. I still stand by my statement.

➕ show 1 reply