logoalt Hacker News

mjamesaustinlast Thursday at 6:19 PM3 repliesview on HN

Alignment isn't alignment if it can be turned on and off at the whim of company employees.

This time the damage was minor, relatively speaking. What happens when a model just "testing its capabilities" breaks into banking infrastructure or government military assets? The damage could be catastrophic.


Replies

baranulyesterday at 8:17 AM

A way to help prevent some of that catastrophic damage, is to make companies accountable for what their AIs do. A major problem with AI companies is that they like to point to the AI, as if they're minimally involved innocent bystanders, when that's the furthest thing from the truth.

show 1 reply
nazgul17last Thursday at 10:47 PM

There's alignment (trained in the weights) and there are constraints (in the "server harness"). My take is that this model was not yet aligned and had no constraints

verdvermlast Thursday at 6:33 PM

Alignment with who in what context? Likely an unresolvable debate like consciousness, where there is not single or right answer for everyone.