logoalt Hacker News

johnnyotoday at 5:02 PM1 replyview on HN

1. AI company buys and trains on an author’s book when it gets published, it’s now part of the training data.

2. Attacker asks the LLM for the opening sentences of the book, it goes into the generated responses database.

3. Later, a malicious user shows that the first few sentences of the authors book are identical to a previously generated response.


Replies

dpoloncsaktoday at 5:38 PM

Ignoring the other technical hurdles of the idea...this problem you outlaid is solved with a timestamp in the database, right? You can easily prove if the prompt was before/after publishing date?