logoalt Hacker News

glubtoday at 7:10 PM0 repliesview on HN

As long as user provides inputs and LLMs stay LLMs, you can waltz through any guardrail. Fable is the extreme case, but it's not that hard if you know what you're doing and know how LLMs and their guardrails work.

Am I saying that guardrails don't work? No, they probably stop a lot of insane people trying insane things. But you don't need Fable-level guardrails to do that. You probably don't even need to do anything during pretraining, or RL, or classification to make sure model refuses to compy with "hack me a bank" or "make me a chemical weapon".

All models will automatically have guardrails just as a result of training on data that gives them intelligence. You have to actually train it to be malicious to produce something what Dario calls "insufficient guardrails".

No guardrail is going to stop a determined person with sufficient intelligence. It only has to stop ones with insufficient one, and even basic guardrail that are just by-product of training is going to achieve that.