logoalt Hacker News

langstoday at 6:41 AM1 replyview on HN

I am working on the same thing right now. However, unlike storing conversations in an external retrieval system, I use a local LLM to store the conversation's KV cache and perform retrieval directly on that cache. The method involves running a prefill pass and, after obtaining the attention scores, filtering for the corpus segments that received attention.

This aligns with the "zero tokens" approach described in this paper. :)

I tested it on the LoCoMo used in this paper, and also LongMemEval, both achieved SOTA results.


Replies

marak830today at 7:35 AM

I was doing something similar where I saved user input/model output in a multi-depth node style storage system (each depth having more precise details) with the focus on the model having accurate user fed information. I was mostly focused on retrieval of accurate / useful information based on user query (injecting the database node as additional, high confidence information)

Once this(Zero-mem) passes it's peer review, I may have to see if my system can handle something similar instead/in addition.

I'm quite excited to see growth in these different ways of eliminating token's.

Long winded aside, @langs, have you published your work on this?

show 1 reply