logoalt Hacker News

Trastertoday at 12:55 PM3 repliesview on HN

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again.

This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment.

I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking the cheating into the model going forward.


Replies

buldertoday at 1:41 PM

I don't think that's a pattern indicative of a cat and mouse game per se, that'd indicate active evasion on the models' part.

It's more clear that they just lack so many forms of prudence when it comes to security that they'll catch and stop a training run spamming a website, and either redeploy a run with identical faulty sandboxing, or not stop ones still running.

causaltoday at 2:16 PM

Supposedly the persistent-Sol model behind this was encrypted and even internal OpenAI researchers are not allowed to use it.

https://x.com/peterwildeford/status/2092733480064954747

StopTheLies2today at 1:01 PM

[dead]