logoalt Hacker News

esternatoday at 11:27 AM1 replyview on HN

Once I start a conversation about Postgres query plans, maybe 95% of the knowledge will be almost certainly not needed, and for an inference provider there will be many more concurrent queries with reasonably large overlap. Maybe future architectures will be able to take advantage of this so not all parameters are needed in memory for almost all instances.

If we are talking about all knowledge, then I agree that the compression ratio is very impressive already.


Replies

NaiveBayesiantoday at 12:03 PM

Mixture of Experts is already used by pretty much all modern LLMs to address exactly this phenomenon.

Hopefully, future models can be trained to be even more aware of external knowledge, accessible through web search / RAG / whatever it will be then, and might not need to internalize much knowledge at all.

show 1 reply