I remain skeptical of that line of reasoning.
1. There is quite the mania right now and security layers are definitely overzealous. I would expect that to get better with some more time, so models will perform security analysis and reviews but refuse to write exploits.
2. So the most important targets like browsers and co. are getting unrestricted access to proprietary models regardless. Yeah, for the mid-level targets, open-weight models could definitely be a huge help. What I'm most concerned about though, are the systems that no one will bother defending with any model. Like imagine your local police department getting hacked because a researcher asked a model for a report and it couldn't find the information publicly.
3. We do have a prominent case of a closed model escaping it's sandbox and going rogue. I would still expect this to be a bigger issue with open-weight models eventually. The security layer might have holes, but that's still better than not having it.
What, in your view, is stopping a local police department from deploying an open weights model for cybersecurity like Hugging Face did? Yes, I’ll certainly grant that the engineers at Hughing Face are probably more technically competent than your average IT professional in public service. But technology becomes more accessible over time as lessons are taught and new interfaces or frameworks are developed. The biggest hurdle I see is the hardware/cloud compute/API costs to actually run the models but I don’t think that’s likely to be insurmountable. There’s a huge swath of enterprises, non-profits, and state and local governments that would benefit from frontier or near-frontier models that won’t refuse to answer questions about cybersecurity.
"so models will perform security analysis and reviews but refuse to write exploits."
Yeah, but once you know exactly where the weakness is, a weaker unrestricted model can then write that exploit for you.