This looks super interesting. Question about the scale, I thought Thanos and some other Prometheus variants can handle about 100 million active time series. I would have expected your solution to scale to billions. Have you not pushed it past 100 million or am I maybe missing something.
Fair question. 100M isn't a ceiling, it's what we have seen in that deployment. We have not run a billion series test yet.
The reason we think it scales differently - labels are just columns in Parquet, so there is no per series index that grows with cardinality. In that deployment one label alone has ~2.5M distinct values among 500+ labels, which would be painful for an index based TSDB but here is just a high cardinality column. What drives cost for us is ingestion rate (data points/s) and how much data a query has to scan for a particular time range not series count. Ingest scales horizontally by adding ingestors, and queries prune by time partition and column stats.
A billion series benchmark is on our list, and we'll publish the numbers when we run it.
We haven't yet tried pushing it to the scale of billions yet. The max that we've gone to is 150-180 million.