logoalt Hacker News

wyrdcurtyesterday at 4:22 PM1 replyview on HN

I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.

Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].

[1] https://transformer-circuits.pub/ [2] https://arxiv.org/pdf/2509.03518

(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)


Replies

utilize1808yesterday at 4:41 PM

Except that the model doesn't hold any malicious intent when doing it. No "deception", no "sabotage".

show 1 reply