logoalt Hacker News

GeneralMayhemtoday at 5:07 PM1 replyview on HN

Yes, that is exactly Dario's concern. Either one of the US labs or one of the Chinese ones will eventually release something with insufficient safety controls for its power level because it gives them slightly better user retention (look how much complaining there is about current frontier models, especially Fable, rejecting requests). Regulation or consortium is how you avoid the prisoner's dilemma.


Replies

glubtoday at 7:10 PM

As long as user provides inputs and LLMs stay LLMs, you can waltz through any guardrail. Fable is the extreme case, but it's not that hard if you know what you're doing and know how LLMs and their guardrails work.

Am I saying that guardrails don't work? No, they probably stop a lot of insane people trying insane things. But you don't need Fable-level guardrails to do that. You probably don't even need to do anything during pretraining, or RL, or classification to make sure model refuses to compy with "hack me a bank" or "make me a chemical weapon".

All models will automatically have guardrails just as a result of training on data that gives them intelligence. You have to actually train it to be malicious to produce something what Dario calls "insufficient guardrails".

No guardrail is going to stop a determined person with sufficient intelligence. It only has to stop ones with insufficient one, and even basic guardrail that are just by-product of training is going to achieve that.