logoalt Hacker News

vessenestoday at 3:12 PM0 repliesview on HN

Anthropic’s Mechinterp did some very fine work on this. TLDR - you can; you train a decoder on neuralese to english and then add a loss function for a roundtrip of english -> neuralese -> english (or possibly n -> e -> n? I don’t recall), giving a pretty strong indication that you have a good ‘translation’.

They published open weights versions of these interpreters for a number of open models sometime in the last year. Very cool idea.

By the way, they concluded CoT often lied, based on the neuralese interpretation.

EDIT: a comment below linked to https://www.anthropic.com/research/natural-language-autoenco..., which is what I was referring to.