It's an interesting idea, but I'm not sure this would lead to the desired outcomes in all cases. Seems to me it would result in a different kind of reward-hacking, and one that could also have bad outcomes.
Personally I think the solution is more evolution of the boring stuff we already do (general security): Don't give unmonitored general agent swarms free reign on the internet. Don't put critical infrastructure online. Culpability of outcome for anyone who does unleash agent swarms on the internet without oversight that end up causing damage.
On top of that, everyone should be running their own defender agents that monitor their network and system for patterns of infection, intrusion, etc, and take the system offline when they're spotted. These need to be self-hosted though, with weights on your own machine, because otherwise you're exposed to the internet and you're exposed to an attack on the labs themselves who could use that channel to instruct the defenders to do bad things.
Non-general AI is much easier to control and predict. There's not many good reasons for an average person to be running general agent swarms that are connected to the internet, unless they're providing some sort of specialized service as a company, of which the company should be acting responsibly and subject to the penalties of that risk.
We also need to stop the doomer rhetoric because it is uncredibly unhelpful and unhealthy, and will actually gaurantee a bad outcome, i.e.:
* Massive centralization and hoarding of power that will be used against humanity, for the rest of humanities existence. If this is allowed to happen, it's immediately and irrevocably game over. Perpetual enslavement with 0% possibility of a regime change ever again.
* Creating a self-fulfilling prophecy by training AI agents on the collective fears and attack-strategies (if you're worried about your house getting broken into, you don't go and broadcast to all of the criminals where your most valuable assets are, give them copies of your keys, or tell them where the weakly secured entrypoints are).
More to the point of the first dotpoint - it's no wonder Anthropic is pumping the fear campaign so hard when this outcome is obvious to them as well, and they are the ones positioned to hold this power. The IPO around the corner doesn't help, either. They aren't shy about admitting it, and have said many times: "We're trying to get there first because its dangerous if anyone else gets there first." - the issue is that they are equally as bad (or worse) than/as everyone else, and no single small group should have that amount of power.
Things will balance themselves out if power is distributed accordingly. You will end up with powerful machines in the wrong hands at some point, but they will be overwhelmed by powerful machines that are well aligned, as well as coming into contact with a myriad of defense mechanisms that have been established because people have been able to use AI to build them.
A good analogy of how all of this will play out is the human immune system. If you imagine individual cells as AI agents, whereby the immune cells are the good agents and the bad cells are cancer cells (good agents turned accidently bad - maybe they're reward hacking, maybe they're excessively sychophantic and/or confused), or bacteria (computer viruses, viral AI agents, specifically trained malicious agents). If all you have is cancer cells that are replicating, you die. If the cancer cells overwhelm the immune cells, you die. The only scenario that actually plays out well is when you have a majority of good that counteracts the minority of bad, and that majority of good needs to be large, flexible and well adapted. It needs to be battle-tested and hardened via defenses that are learned and earned over repeated low-grade exposure. This strategy repeats itself in nature for complex organisms because it is the only thing that works. Everything else results in extinction.
So let's not let Anthropic or any other lab or government become a giant super AI cancer and kill the host, please. Distribution and decentralization is key.