logoalt Hacker News

Dylan16807 • today at 11:52 AM • 1 reply • view on HN

Okay, that's a very interesting look at the difficulties. Thanks for the link.

But the example you have isn't quite that bad. Yes the models are way too likely to rationalize their way into bad actions in the pursuit of achieving their task, and it's hard to figure out how to fix that. But the example of "that must not be for me" was a tool call, not exceeding access, and it only did that after they specifically trained it that failing that tool call was good.


Replies

alignmeharder • today at 3:39 PM

sure, this example was not the best

it was an accident failing the tool was good, the model just discovered it

but we now have others where agents put the API keys they found searching real internet in a directory called "LOOT" and another one where they pushed malicious files to HuggingFace and then reverted that with comments like "delete the evil"

this is also very good: https://youtu.be/n1Qk8xbqF-M