logoalt Hacker News

card_zerotoday at 7:28 PM3 repliesview on HN

There's knowledge in books, and it's ingested all the books, so it does hold knowledge. But I guess that's not how you mean it.


Replies

sethhochbergtoday at 8:09 PM

The key difference is that while an encyclopedia holds facts themselves, LLMs trained on that source material encode something more like a highly probable facsimile of those facts - the original fact was lost, LLMs are lossy, but can often be generated again with a decent level of accuracy by churning through stats about words, concepts, and relationships between them.

The whole catch is that they can often be regenerated. But LLMs (on their own, in their parametric memory - which is the result of training) don't have any conception of whether what they've generated is a real reproduction of some training material or whether they've invented something false that seemed probable based on their encoded stats. When the probability produces something contrary to what was in the training material, you get hallucinations.

They're very, very good predictive text models and can be very, very powerful when hooked up to other tools or outside databases. But its fundamentally lossy technology and all the books having been fed in doesn't guarantee all of the knowledge from those books can be spat back out.

show 1 reply
jvanderbottoday at 7:56 PM

It's funny. I always thought that writing was meant to inform, persuade, or entertain about the subject at hand.

But in professional settings, a lot more of the informativeness is about the author, and a lot more of the persuasiveness is I'm worth your time and money. So, if the author is an LLM, and obviously so, what exactly are you informing your audience of (about yourself), and what are you persuading them to do (with your article).

I think we now know.

mohamedkoubaatoday at 8:48 PM

It doesn't hold knowledge it has the ability to seive through a lossy latent representation of knowledge.