This was a knee jerk response to the first paragraph. They weren't talking about a general crawler, but a subset of Wikipedia, stack overflow, programming docs and github. You can download archives of all of those except github, and github could be queried using the api or GH cli