Hi all - I just wanted to announce a low memory footprint FOSS graph db I've been working on (named Slater, after the Archer character). The most common complaint about Graph/GraphRAG DBs is the cost and memory footprint they consume: many depend on holding the whole dataset fully in RAM, which makes them expensive to run.
Slater starts with the premise of a fixed memory budget applied to an LRU cache of what's on disk, then uses ISAM blocks and DiskANN/Vamana/PQ to allow paging the contents of the graphs and vectors you need into that cache. It's designed around read-heavy-write-light cases, and can be backed either by local disk or by S3/GCS buckets with an optional sized local disk L2 cache as well.
Speaks standard Bolt, and is GDPR friendly (encryption at-rest and in-transit). Multi-user-multi-graph with ACLs, Rust-with-forbid-unsafe and NFS-friendly too (no mmaps). It's Apache licensed.
Please give it a try if you get a chance. Would love any suggestions or feedback.
Full disclosure: yes, it was authored by Claude Code, although I provided the storage model and design it used, along with code samples and influences from other open source projects like FalkorDB and Memgraph for Bolt wire-compatibility.
Thanks
One recommendation: do not let the LLM write your README for you. It can write other docs (though IME they still suck at it), but your README needs to be focused and easily digestible by a human glancing at the project.
LLMs have no concept of focus when it comes to docs, so they spew out way more information than is necessary or helpful to a human reading the document.
Big weakness of Neo4j, etc...?
Since it is a much in demand feature. Why do you think they have not done it themselves already, versus what you have done?
Curious if this is breakthrough? Or there are some negatives that prevent Neo4J from doing it also?
I've wanted to analyze all published scientific papers and authors and their citation relationships in Neo4j but hit resource limits so this is very interesting to me.
Cool stuff!
I still always want to know the numbers per edge. Because there'a a reaon why these graph DBs try to hold the stuff in memory, everything else is slow as heck. We have ~30bytes per edge in our succinct indexes, which makes the 1.5gb wiki dataset clock in at ~50gb.
If I understand correctly you shard the graph and then content adress each chunk:
- How do you decide on the sharding? I'd expect finding good connected subsets with nice memory locality to be extremely difficult computationally.
- How do you canonicalise your graph. Which has also been an extremely difficult problem in RDF land for example. Although that one is at least efficiently solvable.