logoalt Hacker News

wyrdcurttoday at 12:13 AM1 replyview on HN

In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.


Replies

_factortoday at 1:17 AM

It’s not hypothetical. Magic strings are a known and implemented feature for standard model interaction. Nearly impossible to detect unless you know where to look with current technology.

show 2 replies