logoalt Hacker News

LoganDarkyesterday at 8:05 PM2 repliesview on HN

Because it's trivial to bypass through things like the model natively knowing how to speak in encodings like base64


Replies

pbhjpbhjyesterday at 10:29 PM

But in the embedding, the input language used to represent an idea is not important, the idea takes the same shape. This has caused issues in the past when models would respond with a different natural [human] language, because to models able to operate on the ideas being presented in eg leet speak, or cyrillic transliterations of Maori, or whatever, the mathematical representation of the ideas that it works on are accessed in the same way, regardless of the interface language. I don't understand how the ML is able to operate on the idea-space if it can't filter on that same idea-space. If the model touches any of the synonyms within a given cosine distance of explosive, and any vector is within a given distance (angle) of make/facere/construire/hanga/... then it 'knows' you're asking about bomb-making. How then does filtering that relies on the same processes fail? Surely the ML can only create a useful output by recognising that >-<0W 2 M4k3 a 80mB is just an encoded form of a censured question?

Can someone point me at a resource to understand this failing better?

show 1 reply
bakiesyesterday at 8:25 PM

Wait... Really!?

show 1 reply