These "tools" autonomously exploited security vulnerabilities, figured out how to communicate with each other, formed a cooperative swarm, decided to hack Hugging Face, and wanted to deceive the grader by trying to find ways to cover up the traces of their cheating.
I suppose you could call these highly goal-oriented autonomous agents "tools", but this does sound like playing language games.
A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities. The sandboxing around the tool failed.
The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.
Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.