I don't think that 2 is true though. It's exactly the argument we made against open source software.
Yes these models make exploits easier to exploits. Lets use a construction analogy: they've made all of the defects in our buildings easy to see. We have a choice. We either fix those defects, or we ban the tools which lets us see them.
It's clear to me what we do: we use these new tools. Then we fix the defects. Yeah sure it'll mean some work for us but at the end we're in a much much better position.
Anthropic, Open AI and Grok are arguing for hiding the defects. For making us weaker and more vulnerable. For their own profit.
One challenge is that we’re still constantly generating defects even with these tools. And you need the next gen model to spot the mistakes of the previous one. The exploits get more and more complicated and involved of course. But now:
* there’s a long tail of software that just won’t get secured or will take a very long time
* exploits have transformed from an expertise problem to a compute time search.
This is very different than any problem faced before. I’m not saying I’m convinced by the “slow down” approach, but I don’t think it’s as simple as you point out.
Agreed, and if you look at the public information about various attacks the individual vectors are not particurarly sophisticated. The Huggingface incident for example had agents in a sandbox that was about as useful as a wet paper bag, and the attacks on remote infrastructure were basically enabled by bad sanitation. None of the techniques are novel.
These are not hard problems to fix, nor should anyone find it acceptable to have them be so prevalent. A determined human attacker could easily exploit defects like the agents found. Bad input sanitation, SSRF attacks, exploiting stupidly implemented token verification... all techniques that have been widely known for decades and there should be no excuse for publishing software that is riddled with exploits that enable the use of them.