Can someone help me understand the underlying motivation behind this?
It makes sense that some crawlers, in the style of Google, would want to index the entire internet. But what is the point of the same crawler re-fetching a page they already fetched an hour ago? Or possibly all this traffic is just independent entities, each trying to cache the internet? The scale of bot traffic makes this seem unlikely.
What's the motivation behind the same entity re-fetching a page it just fetched less than an hour ago?
Similarly, I've never understood the economics behind the constant rescraping that is flooding the internet or really what's triggering it.
It can't all be agents reacting to user queries. It's confounding how much CPU and bandwidth is getting flushed down the drain.
The most recent data on the internet for advertising, intelligence, etc.
And a lot of bad scrapers.
They are poorly implemented by the "fuck you I got mine" crowd. They will get stuck doing things like trying to run through a calendar that could theoretically go back to the beginning of time and all the way to the end of time. And because that calendar might change, it gets scraped for every inquiry made to the poorly implemented AI system.