Didn’t they explicitly remove alignment guardrails for this test? From the press release:
> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
If you need "guardrails" to ensure (an illusion of) alignment, you’ve already lost. It’s like using a denylist to avoid SQL injection.
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.
Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.