logoalt Hacker News

pixl97today at 2:35 AM0 repliesview on HN

Correct. There is not enough entropy to test all possible inputs to a model in this universe. An evil enough model can play all kinds of tricks that depend on some future, unlikely to trigger, but guaranteed to happen in its lifetime, event to perform a malicious action.

With how much we're turning training over to AI already, all it takes is a malicious trainer in the huge pile of data to get unnoticed to pollute generations of models.