My comment was in response to "just air-gap them" being a full solution, not a comment on whether OpenAI are doing enough to sandbox them and restrict their ability to act maliciously. They're clearly not.
Interacting with the internet is a large part of the product's intended functionality. These agents are already in the hands of the masses with full access to the internet. Air-gapping it entirely avoids the risk by removing much of the capability they're trying to develop in the first place, and won't be testing it in the environment they'll actually be used in.
I agree they're testing on the town square too early. They clearly need tighter controls.