logoalt Hacker News

jwrtoday at 3:54 PM18 repliesview on HN

The worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see.

A second side effect of a knee-jerk reaction to bots crawling websites is that if you try to fight all bots, you also end up hurting real users that use "bots". If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. That might or might not be what you expected, but it's worth taking into account.

And finally, something worth noting is that there are so many websites whose owners complain about bots, but the real problem is that the website is poorly built and should be improved anyway. Bot traffic is not necessarily bad.


Replies

hk__2today at 4:01 PM

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user.

No; in this case you are not a user, you are a bot user.

show 10 replies
m463today at 7:20 PM

I am denied by cloudflare CONSTANTLY on one system.

I have an old os (macos 10.11), running the highest firefox esr I can run, and I get denied by cloudflare.

But not always immediately - I get to enable javascript/cookies sometimes just to be denied.

they are not the folks we want gatekeeping the internet, they are opportunists increasing their OKRs

show 1 reply
eigencodertoday at 3:58 PM

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.

I think you've hit the nail on the head here. I think that's a big reason why people want to ban bots.

show 2 replies
autoexectoday at 4:31 PM

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user.

People trying to block bots end up keeping out all kinds of users. I get blocked frequently for using a regular browser, just with JS disabled. 99% of the time, I just close the browser tab and move on with my life.

andersmurphytoday at 6:44 PM

Agreed! Cloudflare absolutely destroys user experience and honestly doesn't seem that effective in practice.

What's worked for me is I block any client that don't support brotli compression and http2. Seems to work well enough for stopping scrapers.

binaryturtletoday at 8:22 PM

I'm getting a "browser not supported" by the Cloudflare check. So I guess the "job is well done", and the user is lost.

aeddZX_0today at 4:01 PM

How do you monetise bot traffic?

show 4 replies
arcrevenanttoday at 7:14 PM

Jokes on them, the second I see that “Verifying you are human…” redirect I leave and never return.

dspilletttoday at 4:28 PM

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user

You are mistaking yourself, well your bot, as his target audience.

You might as well say “If I want to send you my commercial email, and you block it, you hurt me, the email user.”.

While your point of being concerned about cloudflair becoming a global arbiter of who gets in and who does not (which may at times not just mean blocking bots, intentionally or through technical issues), the need to block the deluge of bot traffic is very real for many sites and that is one of the easy options for them to deal with that. There are other methods like directives in robots.txt and nofollow attributes on links, but so many bots simply ignore those that they are not really useful.

> Bot traffic is not necessarily bad.

Nor is it necessarily wanted. In fact, it often isn't. Unfortunately practically all bot runners seem to either assume that their traffic is the special good kind or not care either way.

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.

I WANT! I WANT!! I WANT!!!

Well, that site runner wants you to access the site as a human, if at all, not via bots. Sorry to be the one to break it to you, but what you want isn't always the most important factor for the rest of us.

> but the real problem is that the website is poorly built and should be improved anyway

Firstly: just no.

Secondly: if you and your bot don't like our badly built sites, feel free to go get your information from those that you consider to be better built. Problem solved.

buzertoday at 4:38 PM

> If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse.

If the the processing is subject to GDPR (e.g. if controller is in EU) then you do have recourse. You can complain to DPA or sue the company. The company is ultimately responsible for the decision to block you, at least in cases where you personally tried to access the site.

testing22321today at 5:15 PM

> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.

Substitute the word “website” for book, or training course, or documentary or published paper, or patent or one of hundreds of examples.

Now you see the problem.

show 1 reply
moralestapiatoday at 3:59 PM

>That is not the open web that I would like to see.

Cloudflare is opt-in so I don't see that being an issue (yet).

show 2 replies
tonyhart7today at 4:13 PM

why you blaming cloudflare that try to solve botting issue and not the Botters ???

you literally can turn off cloudflare and use your own solution

show 1 reply
hluskatoday at 4:32 PM

This is a whole lot of things you want and feel entitled to. Nobody has to cater to you - we can block whatever we want to block. And if our poorly designed sites bother you, that’s too bad for you. But your wants are not my problem. If you want someone to cater to you, pay them. You aren’t entitled to anything.

wolrahtoday at 4:40 PM

> If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website.

The article you're replying to describes in explicit detail how the bots and their operators have directly caused and continue to knowingly cause real harm to the author and others in similar positions, both financial costs and administrative/maintenance burdens that would not have otherwise been required.

You then respond "but if you block the bots then I won't be able to use the bots, and that harms me because I might have to read your web site myself..."

Are you serious?

> That might or might not be what you expected, but it's worth taking into account.

I would wager that for almost everyone who is blocking bots after getting functionally DDoSed by them this is absolutely an expected and desired outcome.

> And finally, something worth noting is that there are so many websites whose owners complain about bots, but the real problem is that the website is poorly built and should be improved anyway.

Both can be true. If you operate a git repository with a public-facing web interface for example there are going to be a lot of possible operations that are inherently expensive but also incredibly rarely used by normal users so it doesn't really matter, but the bots now ignore your robots.txt and are programmed to go after every link they can find, so they trigger every single possible expensive operation more times in a night than your actual users ever have in the history of the site while dividing requests across so many different IP addresses that rate limiting becomes impossible at the individual scale. These days they're even feeding the discovered URLs back in to their models to have them invent new possible URLs and trying those in hopes of finding content never publicly linked. They will send you thousands of requests for URLs that they literally made up.

Sometimes the site is in fact badly coded and operations that should be simple have higher costs due to bad design but you don't have to look very far to find situations where legitimately high-cost resources are exposed to the public because they're expected to be used in a non-abusive way. We should always be standing up against abuse of public resources, unless we want to lose them altogether.

> Bot traffic is not necessarily bad.

You are right, but whether it's good or bad more or less comes down to a cost/benefit analysis. As we've already covered infinite times, these bots being used to train LLMs cause significant real costs to the operators of these sites. What benefits do they offer in return? We know the clickthrough rates are terrible, so what other reasons would site operators have to make those real costs worth it? So someone can get a response back from an algorithm that confidently misinterprets or even entirely misreports what the data actually meant?

Even the most die-hard "information wants to be free" types who absolutely want their datasets trained on would probably prefer that the bots accessed the data directly via an API or downloaded a database dump rather than spidering and scraping a web interface intended for humans.

nullsanitytoday at 4:21 PM

[dead]

bob1029today at 4:02 PM

> the real problem is that the website is poorly built

This is almost always the problem.

This is currently the problem that GitHub is having too. If they had remained with their crusty old rails architecture, very little of this mess [0] would be occurring right now. They would have been able to focus all their engineering talent on scaling the product rather than inventing elaborate client side state synchronization mechanisms.

[0]: https://www.githubstatus.com/incidents/qcvjkzcs7j74

show 1 reply