logoalt Hacker News

RandomLensmanyesterday at 10:06 PM2 repliesview on HN

Which is why with organic intelligence we (sometimes) limit what they can actually do instead of relying on alignment. Can do the same here.


Replies

aesthesiayesterday at 10:16 PM

Absolutely, and we should do that. But it's also directly in tension with getting models to accomplish useful things autonomously. And once you give a sufficiently capable model enough surface area to work with, unless you're able to build a completely unhackable system, any further constraints you put in place are basically advisory. The models in this incident were already sandboxed! Certainly OpenAI's and Hugging Face's security could have been better, but these events point out the risks in relying solely on external constraints on model behavior.

show 1 reply