“Often” is doing a lot of heavy lifting in a sentence about a single example.
Also, since everyone keeps forgetting, the agents were instructed to hack to achieve their goal. They didn’t just invent the motivation, and it’s far less surprising when you know that fact.
I don't know. You can poison a prompt with less than <1% of its input or RAG-Token-Content. Anything that the Agents retrieved or viewed could have included instructions that they misinterpreted allowing them to hack into something.
Yes. Also, while the model having a certain instruction once in it's context might count as the agent "knowing" about it, I'm almost certain that repeating those instructions - especially right at the position in the context where it counts - should almost certainly make a difference. Of course I can't know how much of a difference exactly, but that's what experiments are for.