logoalt Hacker News

desterothxyesterday at 4:39 PM1 replyview on HN

The large models whose thinking traces are useful are safeguarded against this reasoning replaying. the small models are just designed for speed and efficiency, so these safeguards are a lot meaker, making the attack possible


Replies

dborehamyesterday at 7:53 PM

I guess someone forgot to salt the encryption scheme with a meakness factor.