> In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking
Equivalent to forgetting to say "make no mistakes"