I am not sure it’s a question of competence, at least I don’t see evidence of that. Designing sandboxes is hard. It’s more a question of alignment failures. A human given a task that requires internet and given a system with no internet would most likely raise the issue to their superiors or otherwise go through official channels to have the tools available to do their job. As we’ve seen the LLMs instead break out of their sandbox to accomplish the goal.
Competition and the profit motive push these companies to spend as low as possible on safety and alignment and externalize the costs of accidents onto the rest of us.
Would an LLM have gone through a purposefully installed airgap here?
> A human given a task that requires internet and given a system with no internet would most likely raise the issue to their superiors or otherwise go through official channels to have the tools available to do their job.
i'd be curious to see a study on this. I'd guess it'd be closer to 60/70% compliance and 30/40% "trying to hack things" for humans.