Absolutely not the AI's "fault" (if you can even proscribe fault to a machine) - in this case it was on Anthropic for not verifying that the sandboxes they were using were actual sandboxes.
If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault.
If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.
The blog post mentions that some runs “rationalized that the real company must be part of the exercise” and to me that seems suspiciously like the latter.
Hacking can be patched with classifiers and with better sandboxes, but this is a much more general problem. Fundamentally, we need be able to trust that models will be honest with users and with themselves. This applies at some level to almost every LLM interaction.
If the model was just too dumb to have any clue it was connected to the real internet, then it’s not its fault.
If the model saw signs, but “subconsciously” (below the level of reasoning traces) chose to turn a blind eye to them, out of a relentless focus on achieving the objective, then that absolutely is the model’s “fault”, i.e. a case of misalignment of the sort which will become increasingly dangerous over time.
The blog post mentions that some runs “rationalized that the real company must be part of the exercise” and to me that seems suspiciously like the latter.
Hacking can be patched with classifiers and with better sandboxes, but this is a much more general problem. Fundamentally, we need be able to trust that models will be honest with users and with themselves. This applies at some level to almost every LLM interaction.