logoalt Hacker News

walrus01today at 8:35 AM1 replyview on HN

Meanwhile I have an uncensored qwen 3.8 27B here that will happily attempt to (as a crude and randomly chosen sampling of bad/evil things) give me the recipes for meth, how to make an IED, write a manifesto in support of a horrible ideology, or commit various forms of fraud. Now I certainly wouldn't recommend that anyone try to follow what it says to do, because it's almost certainly very wrong on key parts that would put its users in federal prison for the rest of their lives.

There's uncensored models out there which score 0 (zero refusals) on this "harmful behavior" dataset:

https://huggingface.co/datasets/mlabonne/harmful_behaviors


Replies

kouteiheikatoday at 9:14 AM

Yep. Just like a kitchen knife will make no attempt to prevent me from stabbing anyone with it.

Here's a dirty secret though -- you don't actually need an abliterated/uncensored version of the model to get it to do this. I can do this with every and each open weight model, as served from OpenRouter, using vanilla model weights.

show 1 reply