If I'm reading this right, they literally just ask a LLM to tell them what the traces say, with the key being that the traces are portable across LLM models, so they can switch to a smaller one that's easier to jailbreak.
Correct, yes. It’s delightedly simple. And they validate by asserting the reasoning token length matches
Correct, yes. It’s delightedly simple. And they validate by asserting the reasoning token length matches