Why they use different models to decode the reasoning content? Can the the model decode it?
Because stronger models are harder to jailbreak from the paper it says haiku was easily fooled into giving us thinking contents
Because stronger models are harder to jailbreak from the paper it says haiku was easily fooled into giving us thinking contents