logoalt Hacker News

swed420yesterday at 1:35 PM1 replyview on HN

Not in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back.

The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.

The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?


Replies

dparkyesterday at 2:07 PM

> they're using the info to regurgitate in some fashion and serve back.

Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.

> The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?

Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.

They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.

show 2 replies