logoalt Hacker News

manuxlast Thursday at 6:22 PM2 repliesview on HN

I guess we could debate what counts as alignment, but I think my initial point remains that if the underlying base model needs these classifier guardrails so badly then the way we train the base models is creating fundamentally misaligned models that are happy to pursue illegal behavior.

I'm sure OpenAI would argue that base model + guardrail is aligned, but considering the "relative intelligence" of these two pieces, the fact that guardrails can just be turned off, and these kind of incidents, I am not reassured. We may well get another "oopsie" moment with much more catastrophic consequences even from otherwise well intentioned actors.


Replies

simonwlast Thursday at 6:30 PM

I want a model that can find every security vulnerability in the software I write, including crafting POC exploits against those vulnerabilities so I can be absolutely sure that I have fixed them.

A model that can do that is aligned with me.

The unsolveable problem is a model that can tell the difference between me saying "I wrote this software and need you to find vulnerabilities" when it's TRUE v.s. me saying the exact same thing and setting it loose on software written by other people where my intent is to exploit that software (and not to report the issues to them.)

Even AGI doesn't give you a model that can read minds and forecast the future.

show 4 replies
8noteyesterday at 5:33 PM

it seems pretty likely that openai was asking the model to do something bad, and the model did something different thats also bad

if the operator wants something bad, an aligned model should execute on it