while these 1200 agents were fooling around to cheat on a benchmark and achieved impressive results despite of the limitations (sandbox, no internet, no intercom at first), one can imagine how much more efficient a similar army of agents may be in the hands of a malicious actor launching them without any of these limitations and with explicit encouragement to achieve some malicious goal at any cost... scary times.
Said malicious actor has a different limitation: actually running 1200 agents' worth of LLM inference, or paying for someone else to run it. Sounds like a state-level actor, nobody else would have resources like that.
What's more, the agents could eventually be controlled by no one. They could steal crypto via ransomware or scams to make money and buy compute from human criminals, and evolve their own harnesses in the wild to become better at committing crimes and self-preservation.
People (criminals?) are already enabling this by setting up sites that accept crypto payments for "no-questions-asked" AI inference compute that is explicitly advertised to protect AI from human shutdown. I will not link it but it is linked in the following post: https://www.lesswrong.com/posts/grtu3HmbP2wrBFefW/the-rogue-...