an AI alignment researcher said they ran experiments testing exactly this, the model reasoned "this note is not for us. proceed".
Link? Name?
Link? Name?