The authors of the benchmark did not verify that all the tasks were solvable. Apparently, a significant fraction were completely impossible: the given vulnerability could not be turned into a successful exploit.
In hindsight, it seems almost unavoidable that a capable and extremely persistent agent, with lowered guardrails, and faced with an impossible task that it _must_ solve, will start throwing wilder and wilder ideas at it.