LoCoMo is not quite our use case. It tests conversational memory, while Almanac is more for coding agents trying to find their way around a codebase.
These results show output quality at the same token budget, not token savings yet. We’re working on coding agent evals that should be more representative.
We’re working on evals now. We ran a small preliminary LoCoMo test with a 2k-token retrieval budget:
Almanac: 55.7% BM25: 51.8% Supermemory: 47.6% Mem0: 60.6%
LoCoMo is not quite our use case. It tests conversational memory, while Almanac is more for coding agents trying to find their way around a codebase.
These results show output quality at the same token budget, not token savings yet. We’re working on coding agent evals that should be more representative.