I am working on the same thing right now. However, unlike storing conversations in an external retrieval system, I use a local LLM to store the conversation's KV cache and perform retrieval directly on that cache. The method involves running a prefill pass and, after obtaining the attention scores, filtering for the corpus segments that received attention.
This aligns with the "zero tokens" approach described in this paper. :)
I tested it on the LoCoMo used in this paper, and also LongMemEval, both achieved SOTA results.
I was doing something similar where I saved user input/model output in a multi-depth node style storage system (each depth having more precise details) with the focus on the model having accurate user fed information. I was mostly focused on retrieval of accurate / useful information based on user query (injecting the database node as additional, high confidence information)
Once this(Zero-mem) passes it's peer review, I may have to see if my system can handle something similar instead/in addition.
I'm quite excited to see growth in these different ways of eliminating token's.
Long winded aside, @langs, have you published your work on this?