logoalt Hacker News

xpcttoday at 1:50 PM1 replyview on HN

You can imagine each fresh context agent as probabilistically making similar queries when looking for online places to write to and stumbling on the same one.

This becomes even more likely if it's one of the websites that got reinforced during their training process, which they may have used for reward hacking.


Replies

AaronAPUtoday at 2:08 PM

I wonder if the sort of algorithm which would break this sort of swarm alignment would also break watermarking.