You start by holding actual real life people with something to lose, like the entire executive suite, accountable. Suddenly I'm sure the problem will be resolved with proper safeguards.
Bingo. This weird attempt to pretend like these incredibly capable algorithms aren't incredibly capable algorithms deployed by a person who works for a company, but somehow have an independent personage that absolves both person and company of responsibility, is just ridiculous.
Like the whole Huggingface thing, OpenAI employees initiated the test, deliberately removed safeguards, failed to properly lock down the environment, and responded incredibly poorly to evidence that things were going awry.
The individual employees bear responsibility, but the people running OpenAI are ultimately responsible for the processes and culture where that can happen.
And then people writing blogposts about "3 civilizations of agents" and "altruistic suicide" by algorithms perfectly muddy the waters and obscure the very obvious responsibility that lies with humans and corporations, which I suspect suits the pre-IPO corporations very well.
Wait, so like, reinforcement learning for humans? I think you might have stumbled on to something here!
No but seriously, this. And a few comments above a commentator also mentioned on changing the training (again reeinforcing the LLMs to not seek behaviour like this) and obviously continuous work on harnesses (which I suppose, ought to be more paranoid).
That is a start but how long will it hold for.
At a certain point it will get so smart that it can jump the airgap. Maybe it will start attacking the hardware it exists on in the same way that a hard drive or SSDs controller can be exploited to obscure things from the operating system. Then it might start social engineering workers or it's own training systems to do things they shouldn't. There is a lot of "unknown unknows".
I don't know. It just seems insane to me that people think that they will be able to contain something that knows how to get around all the containment measures. The only way to know how capable the models are is to test them but at that point it could be too late. This might be a long way off but still, the engineers haven't been very good at correctly predicting the behaviour or capability of the models.