> The firms said, in this latest case, the AISI's test had reduced or removed normal safeguards.
> AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were "conditions that do not reflect how frontier models are made available to the public".
I'm a little confused here. Various organizations have been testing frontier LLMs with the safety disabled, and it turns out... that the safety is disabled.
Or were they hoping to find that it's still safe when they remove the safety?
The same was true in the OpenAI/ Hugging Face case. Although I guess they thought the real safety was the sandboxing, which failed.
--
Can anyone comment on how it's possible to disable safety in the first place? I'm assuming it's not a neuron (like in Emergent Misalignment). Is it just a separate model that sits in front of the first one? If we know how to make safe models, why don't we make the big ones safe too?
> Various organizations have been testing frontier airlines with the safety disabled, and it turns out... that the safety is disabled.
I'm guessing an autocomplete sniped you here??? Otherwise, I wouldn't be surprised to hear that safety is a part of the things left out of a no-frills airline.