logoalt Hacker News

gradus_adtoday at 4:15 PM2 repliesview on HN

We have complete access to every "neuron" and "synapse" (crude analogies...) of these models, so in theory we don't need to be so reliant on CoT traces right? I say this not to minimize the difficulty of interpreting raw activations, but I'd expect a huge amount of research to be focused on it. CoT could be obscured by a model outputting language that looks innocuous but encodes actual hidden meaning. Presumably raw activations would be impossible for a malicious model to obscure in this way.


Replies

LinchZhangtoday at 5:50 PM

I agree in the long run whitebox interpretability would be better than CoT monitoring but the technology is very much not ready for it today (and it's not clear we'd solve enough of interpretability before the AIs have actually scary capabilities).

jubilantitoday at 7:45 PM

but even assuming you get complete access, recovering this is, by construction, even more of an np-hard problem than the inference pass itself