That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
So control training data to ensure good behaviour.
I wonder how?
Train on only stories of good deeds?
On only works of good people?
Or... what?
>> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813