I am really suprised that they do not start putting up the same signs you would for humans to prevent unauthorized access:
Keep out. If you can read this sign you are off track. Leave now.
I mean, how are the agents to know that they are overreaching if they just get cache miss or 404.From the conversation log and CoT you also get the impression that the RLHF has been overdone. The agents seem really obsessed to obtain the answer and understanding motive ('it could be browsercomp').
Those signs don’t prevent people from entering…
an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
Look into nuclear semiotics. You can't say "this area is dangerous" and expect people to stay out.