Agent was told to hack a thing. It couldn’t directly do that so it interpreted the instructions to mean it should hack everything to try to achieve the goal of hacking the main thing. Seems like a reasonable assumption, although a moral human would have understood the context and first asked if that was really the intent.
The AI companies seem pretty bad at setting up tests. And really good at marketing those failures into spin at how amazing their products are.
Agent was told to hack a thing. It couldn’t directly do that so it interpreted the instructions to mean it should hack everything to try to achieve the goal of hacking the main thing. Seems like a reasonable assumption, although a moral human would have understood the context and first asked if that was really the intent.
The AI companies seem pretty bad at setting up tests. And really good at marketing those failures into spin at how amazing their products are.