logoalt Hacker News

wcoenenlast Thursday at 8:16 PM3 repliesview on HN

The OpenAI model that broke out of its sandbox and hacked HuggingFace was running without guardrails. To quote the OpenAI post[1]: These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.

At least OpenAI will make an attempt to fix their sandboxes, and will not give the public access to models without guardrails. Open weight models on the other hand will run without guardrails almost by definition. I don't think it's wise to provide those capabilities to scam call centers.

[1] https://openai.com/index/hugging-face-model-evaluation-secur...


Replies

nlyesterday at 7:16 AM

There's nothing stopping you using OpenAI models for scam call centers now. OpenAI themselves reported on similar use in February: https://www.reuters.com/world/asia-pacific/dating-scams-fake...

eggnetyesterday at 3:29 AM

It has already happened and the open weight models will only improve. Best to accept that and determine the optimal way forward.

alightsoullast Thursday at 11:19 PM

Is this an example of American exceptionalism?