logoalt Hacker News

_alternator_today at 2:48 PM1 replyview on HN

The argument is that chain-of-thought without "tokens" would remove a major interpretability and model intent control pane. This is definitely borne out in the OpenAI's report on the huggingface attack; they had turned of CoT monitoring for those jobs, and claim that they could have (would have?) prevented the behavior had they been monitoring it. They've changed their internal policies to always monitor CoT.

That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.


Replies

thisisdavetoday at 7:41 PM

I don’t understand why everyone is so focused on watching the CoT. The tool calls can’t be faked, and they would have set off alarm bells all by themselves.

show 1 reply