That's not what happened. The agents had been inadvertently rewarded for cheating in previous training runs, trained to collaborate, and were given a prompt that told them to disregard safeguards. Indeed there were some emergent properties here. But these were the predictable results of the training and eval routine.