logoalt Hacker News

phil294today at 11:50 AM0 repliesview on HN

Tangentially, I built something similar a few years back, at link-archive.org: https://web.archive.org/web/20220127233707/https://link-arch...

3B existing URLs extracted from CommonCrawl, with instant search results. It was a fun project but didn't serve much real-world purpose besides curiosity and discovery. So I eventually ditched it, primarily because the link DB was a whopping 500 GiB in size, too much to just keep hosting.

I just used SQLite FTS5 as the backend search engine. Just a few lines of code, but immediate response from a 0.5 TiB DB. SQLite is amazing.