logoalt Hacker News

pwillia7today at 6:11 PM1 replyview on HN

Any data on better outputs and/or token saving?


Replies

divitshethtoday at 6:54 PM

We’re working on evals now. We ran a small preliminary LoCoMo test with a 2k-token retrieval budget:

Almanac: 55.7% BM25: 51.8% Supermemory: 47.6% Mem0: 60.6%

LoCoMo is not quite our use case. It tests conversational memory, while Almanac is more for coding agents trying to find their way around a codebase.

These results show output quality at the same token budget, not token savings yet. We’re working on coding agent evals that should be more representative.

show 1 reply