logoalt Hacker News

Jskeweltoday at 6:23 AM3 repliesview on HN

Impossible. The majority of websites firewall automated crawler traffic (because of the rise of the bots), only making exceptions for the largest search engines. There is no possibility of starting a new crawler.


Replies

Cakez0rtoday at 6:59 AM

It would be interesting to see a decentralised, residential collective that builds and publishes an index. There are surely enough interested people on HN alone that would be willing to run software at home to scrape a small slice of the internet.

show 1 reply
mitxelatoday at 7:04 AM

The majority of websites try to do that but they do not catch as much traffic as they think they do. A starting point for a scraper is to run it on your home connection in an undetectable web driver framework such as zendriver.

dwedgetoday at 9:10 AM

This was a knee jerk response to the first paragraph. They weren't talking about a general crawler, but a subset of Wikipedia, stack overflow, programming docs and github. You can download archives of all of those except github, and github could be queried using the api or GH cli