logoalt Hacker News

limecherrysodatoday at 6:01 PM0 repliesview on HN

Gemini uses MoE and context caching, which is a similar approach.

You are not really accessing the biggest frontier model every time, and you're not really doing an end-to-end LLM request on each prompt.

I would go so far to say frontier models have peaked and improvements from here come from clever (or very elaborate) harnessing. "LLLMHs" - Large Large Language Model Harnessing !