Here's my impressions of your algorithm:
1. read each site
2. rent a 4090 with https://vast.ai to run vllm
3. let llm model invent its own category and tag names freely
4. save 1KB of metadata each
a. a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags.
5. `code is going up as open source` soon (TM)
The technical details are on another page: https://alexmorleyfinch.github.io/marlin/history/v1/article/...
Your impressions seem about right, but there are a few control steps it seems.