logoalt Hacker News

heaney-555today at 8:55 AM2 repliesview on HN

>and (2) are prompted to hack

Sure but the problem in the HuggingFace incident is that they were not.

>You cannot prevent (2) via any alignment process

Of course you can. Go ask Claude Fable to create a malicious virus and it'll refuse.

>Just remove hacking data from the training dataset and you're done.

That's not how this works. The same skills that allow for debugging and writing safe code can also be used to hack.

https://en.wikipedia.org/wiki/Dual-use_technology


Replies

seba_dos1today at 10:28 AM

> Sure but the problem in the HuggingFace incident is that they were not.

Of course they were, even if indirectly.

cyanydeeztoday at 9:02 AM

It is amusing that to "align" a LLM, first you must give it all the things "not to do" and the "not" part is clearly easily lost and you must constantly inject that into their context when it's clearly that they wouldn't hack if they couldn't hack and their intent wasn't given as "hack this".

The openai rogue hacking, if performed by a nation state, would seriously be taken with stern words and likely sanctions depending on the relationship between the two states.

But instead it's treated like a marketing stunt by all liable parties.

show 2 replies