logoalt Hacker News

pastoday at 8:52 AM1 replyview on HN

Arguably the question is whether it's economically feasible to self-host something similar.

Are you self-hosting Google or Bing? No, but we have quite a huge ecosystem of full-text search tools with PageRank, with options to scale to almost Google scale (if you have the money). After all LLM training starts with the same crawl mechanism.

As long as barriers to entry is not too high (ie. it makes sense to take the risk to start a business that provides something similar - usually for a niche) market forces work.

We have the classic empirical chart reproducing microeconomics.

https://www.fda.gov/about-fda/center-drug-evaluation-and-res...

And setting up a pharma plant is also very capital intensive.

Here the obvious barrier to entry is completely artificial. (Which provides an incentive to spend a lot of money on R&D -- though it naturally raises the question of Pareto efficiency.)


Replies

spwa4today at 9:24 AM

You're missing the point, what will happen is this:

1) in things like tax law, registering with city hall, dealings with the DMV, your phone subscription, insurance contract, ... you will find that one of the new fine prints in the contract will be that you're not allowed to use AI to communicate with them.

2) because of how SynthID works (you need the SynthID keys to verify, which are secret. So the only way to find if text is ChatGPT/Google/Anthropic watermarked is to ask ChatGPT/Google/Anthropic), government and large companies can enforce this against you. That is what the watermark is for. To end any insurance claim written by AI with "you're not allowed to submit AI written insurance claims" and refuse it outright there and then.

"Sorry your request was AI watermarked and pursuant to law 234 of 2025/03/11 chapter 3258 paragraph 33 decile 1299 we hereby close it without response"

3) when they reply, however, they use a custom model that also has custom SynthID keys. You will not even be able to tell their responses are AI written, or at least, you won't be able to prove it. You won't be able to enforce any AI-related rights (ie. the right to talk to a human) you have under the law against large companies.

In other words: this is to make sure that all the advantages AI provides are available to deny your unemployment claim, and to Verizon to charge you more, but completely inaccessible TO YOU when you want to change to a cheaper subscription. They can inundate YOU with AI-written requests BUT YOU CAN'T.

Self-hosting helps because it prevents them from verifying if your responses are AI written, because you can generate non-watermarked AI text and so there is a level playing field.

show 2 replies