logoalt Hacker News

dns_snek • today at 5:27 AM • 2 replies • view on HN

I think you have that the wrong way around. Archive services have to explicitly bypass paywalls so most local newspapers are simply unsupported.


Replies

slow_typist • today at 7:26 AM

I elaborate: the local newspaper was perfectly archivable for years. Then it stopped.

The CMS usually allows search engines to get all content. The feature is to know human browser operators and make them pay, or know archiving/unblocking services and block them.

walrus01 • today at 6:10 AM

Based on some of my recent experience with building a scraper for purposes that aren't a news website, what has changed in the last 12 months or so is the availability of very low cost (or free, if you can run it on a 256 or 512GB system in your own office) LLM that are good enough to give it a target of something and have it build a custom scraper/paywall bypass profile on a per site basis. I'm referring specifically to things that can be run in harnesses and score well on terminalbench 4.0 and SWE LLM benchmarks.

When previously nobody would have gone through the effort to build and maintain a custom scraper against a moving target, for some small to medium size city's newspaper, now it's just one of a myriad of 'scraper profiles' you can have a near fully automated tool create.