logoalt Hacker News

dborehamyesterday at 4:12 PM1 replyview on HN

Hmm, ok. So the attack doesn't involve decrypting the payload, only getting the server to do so. Since a model will do that if you just ask, what's so special about the attack?


Replies

desterothxyesterday at 4:39 PM

The large models whose thinking traces are useful are safeguarded against this reasoning replaying. the small models are just designed for speed and efficiency, so these safeguards are a lot meaker, making the attack possible

show 1 reply