logoalt Hacker News

keedatoday at 2:01 AM0 repliesview on HN

A more important quote from that link:

> With our specific experiments, we find that the largest LLMs don’t memorize most books–either in whole or in part.

This technique has only ever been made to work with a vanishingly small number of extremely popular works, probably because they are so overrepresented in the training data set.

The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.