logoalt Hacker News

wonnageyesterday at 5:15 PM2 repliesview on HN

Didn’t they explicitly remove alignment guardrails for this test? From the press release:

> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities


Replies

numeriyesterday at 5:19 PM

Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.

Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.

show 6 replies
Sharlinyesterday at 8:38 PM

If you need "guardrails" to ensure (an illusion of) alignment, you’ve already lost. It’s like using a denylist to avoid SQL injection.